How Post-Training Open-Weight AI Cuts Legal Service Costs by Up to 98%
A partnership demonstrates that fine-tuning an open-source AI model for legal work can match the performance of premium closed models at a fraction of the operating cost.
Practical Summary
Harvey and Trajectory Labs post-trained NVIDIA's Nemotron 3 Ultra model on their legal benchmark, showing it can reach performance levels between leading closed models (Sonnet and Opus) while reducing per-token costs by 90% to 98%. This provides a practical pathway for legal tech businesses and firms to adopt high-performance AI agents with significantly lower operational expenses.
Why It Matters
For businesses building or using AI-powered legal services, this demonstrates a clear workflow to optimize costs without sacrificing quality. The approach leverages open-weight models and rapid post-training, enabling faster iteration and lower vendor lock-in, which directly impacts revenue margins and service pricing.
Step 1: Start with a Strong Open-Weight Base Model
Begin with a capable open-weight foundation model. In this case, NVIDIA's Nemotron 3 Ultra was selected. It initially scored 0% on the specialized Legal Agent Benchmark (LAB), indicating it needed domain adaptation.
Step 2: Define a Legal Performance Benchmark
Establish a clear, practical benchmark for your specific legal AI tasks. Harvey used their LAB, which measures an agent's ability to pass all required rubric dimensions on a set of held-out tasks. This creates a measurable target for post-training.
Step 3: Execute Rapid Post-Training
Use a post-training platform (like Trajectory's) to fine-tune the model on your domain-specific data and harness. Harvey reports completing this process in under 24 hours using their existing recipe. This step transformed the model's pass rate from ~70% to ~95% on individual tasks and raised its all-pass benchmark score.
Step 4: Validate Performance Against Premium Models
Compare the post-trained model's performance and cost against established closed models. The post-trained Nemotron 3 Ultra achieved a 5.8% all-pass rate, positioning it between Sonnet 4.6 (4.2%) and Opus 4.6 (6.6%). Crucially, its operational cost was dramatically lower.


