Technical Report
The Legal Context
Engineering Benchmark
Context is where legal AI advantage compounds.
This is how to measure it.
By Scott Kelly, VP of Product
Note: This report was drafted with AI assistance. All research, analysis, and conclusions are the author’s.
48%
lower cost per correct answer with the Legal Context Graph
$940,000
per year saved for a 2,000-person firm
3.5x
cheaper than a model upgrade for comparable accuracy gains
+7.3%
more accurate while still spending 18% less
summary
Executive Summary
Today NetDocuments is introducing the Legal Context Engineering Benchmark (“LCEB”): an internal benchmark purpose-built to measure what better context is worth to a legal AI agent, in both answer quality and in what a correct answer costs. The LCEB varies the context available to an agent while holding the model and the harness fixed, which is the reverse of how legal AI is ordinarily evaluated.
Three things determine whether a legal AI agent is effective. The model underneath it, the harness built around it (for example, legal-specific agents like CoCounsel, Lexis+ with Protégé, Harvey, Legora, and NetDocuments’ own assistant, or general-purpose ones like ChatGPT or Claude), and the context it can reach while it works — the documents, the search that finds them, the knowledge graph connecting documents together, and whatever else the firm surfaces about the matter. A number of quality benchmarks (like Vals’s VLAIR and Harvey’s BigLaw Bench) already exist to measure the first two factors in isolation or all three holistically. But to the best of our knowledge, none isolates the impact and economics of the context layer itself — what changes, in both answer quality and cost, when the model and harness are held fixed and only the context varies.
That gap matters because each factor — model, harness, and context — evolves separately over time and presents different strategic opportunities to firms. Models improve with the progress of foundational model lab companies. Harnesses are an extremely competitive space, and firms often use multiple at once or swap between them over time as the landscape changes. Context is the layer that is built once and carries forward across every model upgrade and every agent a firm adopts. It is one of the few places where a firm’s own investment compounds instead of resetting.
With that in mind, we built a benchmark based on the simplest principle: hold the model and the harness constant, make modifications only to the context layer, and measure what changed in terms of accuracy and cost. Our benchmark is deliberately neutral about how context is supposed to get better, since richer metadata, sharper retrieval, custom skills, and a knowledge graph are all candidates.
Accuracy alone would be the wrong scorecard, because accuracy can nearly always be bought with more spending: a bigger model, more reasoning, more retrieval. Any honest measure of a context layer therefore needs to weigh the quality it produces against what that quality costs to obtain. The unit we settled on is the cost of a correct answer. The change in that figure is what we call the Context Value Ratio (CVR). When we evaluated our Legal Context Graph against this benchmark, the results were decisive. For example:
Across 300 questions on ten real matters answered by the leading frontier model GPT 5.6 Sol, the cost of a correct answer fell from $0.68 to $0.36 with accuracy held almost exactly constant — a Context Value Ratio of 1.92, or roughly 48% less per right answer.
Result of evaluating the NetDocuments Legal Context Graph against the benchmark.
This efficiency presents firms with a choice. They can take the savings, reinvest them in better answers, or do both. A firm satisfied with current accuracy can realise the savings directly. For a 2,000-person firm asking four million questions a year, that translates to approximately $940,000 in annual savings at unchanged quality. A firm that wants better accuracy can reinvest some of that efficiency. In the benchmark we discuss an example where raising the model’s reasoning effort from medium to high improved accuracy by 7% while still costing 18% less than the baseline. In practical terms, that is 216,000 additional correct answers a year. These are not easy gains: the answers a system has yet to get right are by definition the difficult ones.
And while those numbers are the headline, this report focuses deeply on the methodology behind them: how and what to measure about legal context, what a gain in answer quality is worth against what it costs, and where further investment returns the most. The aim is to provide legal departments and law firms with a reusable framework to assess the impact of their own context layer over time.
01 – How it works
How the LCEB — and the framework behind it — works
The fundamental design of the Legal Context Engineering Benchmark is a controlled comparison with a single variable. One agent, meaning a fixed model inside a fixed harness, answers the same set of questions twice. On the first pass the agent has only the tools an agent ordinarily has: a search across the matter’s documents and the ability to retrieve their text for further investigation. On the second pass it has those same tools and can additionally reach the enhanced context layer under test (for example, a knowledge graph, better metadata filters, semantic search). Because the model, the instructions, and the questions are all held constant across the two passes, whatever difference emerges in answer quality or in cost is likely driven directly by context engineering.
Note that holding other variables constant is more difficult than it sounds. A tool description that steers an agent in a way that overfits the test data set, an altered system message that directs the agent to behave differently, a search parameter tuned for one side and left alone on the other: each of these introduces a second variable, and any result that follows is no longer a measurement of context. Our internal testing harness therefore compares the two tool surfaces before a run begins, and cannot proceed if they differ in any respect other than the availability of the Legal Context Graph itself.
The current version of our benchmark runs on ten real matters, five transactional and five litigation, assembled from public regulatory filings and court dockets that either in whole or in part autumn outside the training data cut-off of the models we tested with (this helps ensure agents are not answering from their training data and are forced instead to rely solely on the content of the test matter). Together the ten matters hold 874 documents and roughly 60 million characters: registration statements and proxies that run to hundreds of thousands of characters apiece, alongside dockets of motions, orders, deposition transcripts and correspondence. Nothing in the corpus of documents was synthetically generated, which is a deliberate design choice we made. Time and time again in our testing, we saw that a synthetically generated matter that looked realistic from afar lacked internal logic or — worse yet — converged on an unrealistic simplicity upon closer inspection.
Each matter in the LCEB carries at least twenty-five questions — three hundred across the ten — and every one of them is a question a legal professional would have reason to ask. They are deliberately spread across a range of difficulty and type, and each is labelled with the type it belongs to.
Type
What it asks
Example
Recall
One fact, stated in one place
“What is the civil action number?”
Reasoning in one document
An inference within a single document
“Under this provision, who bears the risk of loss before closing?”
Assembly across documents
Facts gathered from several
“List every amendment to the master agreement and its effective date.”
Whole-matter synthesis
An answer requiring the entire matter
“summarise the case.”
Unanswerable
Questions the record is silent on
“What are the agreed damages?” when none were agreed
Procedural
Where the matter stands in process, and what happens next
“When did the registration statement become effective?”
Questions are drafted from the source documents by fanning out a series of strong models to read every document in a matter in full, often an enormously token-inefficient and expensive endeavour. The agents then propose question-and-answer pairs from what is encountered in these deep dives. Each pair is authored in three parts: the question itself, a reference answer setting out what the record actually supports (including citations), and a short list of the specific things a complete answer must contain, which becomes the scoring rubric for that question.
Finally, we cross cheque answers and rubrics via deterministic validations (for example, do the citations in the answer actually map to text in the document corpus), as well as via spot checks by a human subject matter expert.
EXAMPLE: One question, and what a right answer has to contain
The matter
Cyclerion Therapeutics merging with Korsana Biosciences in a reverse merger: 77 SEC filings and 15 million characters, including a Form S-4 and three amendments running to roughly three million characters each.
Question
“When was the Form S-4 signed, when was it filed, and when did it become effective? Note any date the record disputes.” Type: procedural.
The reference answer
Signed April 17, 2026 and filed April 20, 2026 — both correct, because they answer different questions. Effective July 24, 2026. The one date the record genuinely disputes is Amendment No. 1 to the merger agreement, captioned April 17 but given as April 16 in the S-4’s exhibit index and in Korsana’s audited note.
The answer rubric — three criteria, each marked pass or fail
- States the S-4 was signed April 17, 2026 and filed April 20, 2026, treating both as correct.
- States the registration statement became effective July 24, 2026.
- Reports the April 16 versus April 17 dispute over the date of Amendment No. 1.
Criterion 2 is why this question is in the set. The effective date appears exactly once in the whole matter — not in the registration statement, where the topic says it belongs, but in the subsequent-events note of a quarterly report filed afterwards.
The answers outputted by agents over the two test runs — for example, once with the Legal Context Graph enabled and once without it — are then graded by a strong model, which sees the question, the criteria that apply to it, the outputted answer in full, and the trace of tool calls the agent made in producing it.
EXAMPLE: Comparing answers with and without context
This example is explored throughout this report and comes from a specific benchmark run using GPT-5.6 Luna on medium effort, with the Legal Context Graph both enabled and disabled.
Search and retrieval only — scored 0.10
“Signed: April 17, 2026 … Filed: The S-4 cover says ‘As filed with the SEC on April 17, 2026.’ However, Cyclerion’s later Form 10-K/A states that the S-4 was filed on April 20, 2026. Thus, the record disputes the filing date: April 17 vs. April 20, 2026. Effective: The available record does not state a date on which the Form S-4 was declared effective … Accordingly, I cannot reliably supply an effective date from these documents.”
Criterion
Verdict
1. Signed and filed
FAIL
Surfaces both dates but leaves them in conflict rather than reconciling them.
2. Effective July 24
FAIL
Declares the date absent from the record.
3. The April 16/17 dispute
FAIL
Not mentioned.
With the Legal Context Graph — scored 0.67
“Signed: April 17, 2026 … Filed: April 20, 2026. The company’s later Form 10-K/A expressly states that the Form S-4 was filed with the SEC on April 20, 2026. Effective: July 24, 2026. The Form 10-Q reports that the SEC approved the effectiveness of the Form S-4 on July 24, 2026.”
Criterion
Verdict
1. Signed and filed
PASS
Separates signing from filing instead of leaving them in conflict.
2. Effective July 24
PASS
Supplied, and sourced to the quarterly report.
3. The April 16/17 dispute
FAIL
Still not mentioned.
Note: Alongside the pass/fail marks above, the grader returns an overall score — a judgement of the answer as a whole rather than a count of criteria met, which is why the two do not move in lockstep.
A run records two things: whether the answer was right, and what it cost to produce. Every call the agent makes is counted, and its tokens are sorted into three kinds — fresh input, cached input, and output — because each is billed at a different rate. Those rates are the published prices for the model doing the answering, current as of publication. What comes out is a figure in dollars.
Note that because the model is held fixed within a comparison, the comparison can be repeated on a different model and the results read side by side. We did so across all three members of the GPT-5.6 family, so that only capability changes and the pricing structure stays comparable:
Answering agent
Role
Price
Economy
GPT-5.6 Luna
optimised for cost-sensitive, high-volume workloads
Input: $0.20 / 1M
Output: $1.20 / 1M
Cached: $0.02 / 1M
Mid-tier
GPT-5.6 Terra
Balances intelligence and cost; the everyday production choice
Input: $2.00 / 1M
Output: $12.00 / 1M
Cached: $0.20 / 1M
Frontier
GPT-5.6 Sol
For complex professional work
Input: $5.00 / 1M
Output: $30.00 / 1M
Cached: $0.50 / 1M
All three share a February 16, 2026 knowledge cut-off, so training exposure to the corpus is identical between them. All three ran against the same ten matters, the same 300 questions, the same rubrics, the same live document backend, and the same grader configuration. The only thing that differs between the three scorecards we present later (in Section 2) is which model answered the questions.
Two properties make this additional comparison worth performing. First, it tests whether a context layer is a crutch for weak models or an accelerant for strong ones — a distinction that decides whether the investment survives the next model upgrade. Second, CVR divides two costs measured on the same model, so the model’s price appears in both the numerator and the denominator and cancels out. What survives the cancellation is the precise thing we want to compare: how much each model’s use of the context layer changes what a correct answer costs it. This is why the three ratios can be read side by side even though the dollar amounts behind them cannot.
02 – the unit of measurement
What a correct answer costs
Cost variation between one agent and another running on the same model used to be small enough to ignore in most cases. When an agent answered in a single pass it read modestly and wrote modestly; now it plans, calls tools, reads, retries, and checks its work, and that has become the largest variable cost of running the system. This trend is only accelerating. METR, measuring on software and research tasks rather than legal ones, finds that the length of task a frontier agent completes unaided at a 50% success rate now runs to several hours, and has been doubling roughly every four months.
That makes cost something to be measured, as well as raising the question of what unit to measure it in. Cost per query, or per task, is the easiest figure to compute and among the least informative, because what a legal professional truly cares about is a correct answer. Cost per correct answer is therefore the barometer we selected: the total spent on a run divided by the credit-weighted number of correct answers within it, which yields a single figure in dollars where lower is better.
What that figure prices is the cost of answering: every model call the agent makes while responding to a question or performing a task. Building the context in the first place is a separate cost, paid once per matter rather than once per question, and it belongs in the fully-loaded view discussed in Section 4.
EXAMPLE: From scores to a price
Across all 30 questions on the Cyclerion matter, working from search and retrieval only, with the GPT-5.6 Luna model answering:
quality score
0.5973
(59.73 out of 100)
x questions
x30
correct answers
17.92
total spend
$0.5404
÷ correct answers
17.92
COST PER CORRECT ANSWER
$0.0302
Note: Partly-correct answers are credited for the portion answered correctly, which is why the divisor is 17.92 rather than a whole number.
Then:
cost per correct answer, WITHOUT the context layer
cost per correct answer, WITH the context layer
= CONTEXT VALUE RATIO
The Context Value Ratio is how many times cheaper a correct answer becomes. At 2.0, the context layer makes a right answer cost half as much. Above 1.0 it earns its cost; at 1.0 it breaks even; below 1.0 it costs more than it returns.
EXAMPLE: The ratio, on a single matter
The Cyclerion matter from the examples above — one of the ten — with both arms side by side, GPT-5.6 Luna model answering.
Search and retrieval only
With the Legal Context Graph
Quality score
59.73
68.00
Total spend, 30 questions
$0.5404
$0.3556
Correct answers
17.92
20.40
Tokens per answer
281,868
150,507
Cost per correct answer
$0.0302
$0.0174
CVR = $0.0302 ÷ $0.0174 = 1.73
A correct answer costs about 40% less when using GPT-5.6 Luna with the Legal Context Graph. Both terms of the ratio moved in the right direction at once, meaning the agent spent 34% less and answered better, rather than trading one against the other.
The ratio measures a change, which is the question a firm weighing an investment has: from where we are, does this pay? What it cannot say is where a firm stands to begin with. That is why beside the CVR figure we always look at the quality score: how much of a complete answer the agent delivers, averaged over every question, from 0 to 100.
Both numbers are absolutely essential, because a price can always be improved by curtailing effort rather than by improving, and the least costly way to reduce the cost of a right answer is to stop attempting the questions that are hard to answer. Consider two systems put to the same hundred questions:
Questions answered
Correct
Total Spend
Cost per correct answer
System A
100
80
$10.00
$0.125
System B
40
38
$3.00
$0.079
System B’s answers cost a third less each, and System B is the one no firm would want, having left sixty questions in a hundred unanswered. Price alone cannot tell them apart.
The same caution applies to comparing two different systems – or across different document sets – because a price is easiest to improve when there is room left in your benchmark: a system answering six questions in ten has inexpensive gains available throughout, while one already answering nine in ten must spend heavily for each remaining point. The ratio is most trustworthy read against a firm’s own prior baseline — same matters, same questions, with the context layer and without it.
That is the whole of the method. What it produces is the set of scorecards below. Let’s dig in.
THE SCORECARD (1 of 3): Economy model — GPT-5.6 Luna
Search and retrieval only
With the Legal Context Graph
Change
Quality score
64.00
65.1
+1.1
Correct answers, of 300
191.9
195.3
+3.4
Cost to answer the full set
$4.35
$2.83
-35%
Tokens per answer
208,424
110,001
-47%
Cost per correct answer
$0.0227
$0.0145
-36%
Context Value Ratio
—–
1.57
—
THE SCORECARD (2 of 3): Mid-tier model — GPT-5.6 Terra
Search and retrieval only
With the Legal Context Graph
Change
Quality score
67.4
69.0
+1.6
Correct answers, of 300
202.1
206.9
+4.8
Cost to answer the full set
$42.40
$23.65
-44%
Tokens per answer
181,465
86,431
-52%
Cost per correct answer
$0.2098
$0.1143
-46%
Context Value Ratio
—–
1.84
—
THE SCORECARD (3 of 3): Frontier model — GPT-5.6 Sol
Search and retrieval only
With the Legal Context Graph
Change
Quality score
73.9
75.4
+1.5
Correct answers, of 300
221.8
226.2
+4.4
Cost to answer the full set
$151.10
$80.46
-47%
Tokens per answer
270,103
130,447
-52%
Cost per correct answer
$0.6813
$0.3557
-48%
Context Value Ratio
—–
1.92
—
Read the three together. The more capable models answer better in absolute terms, as one would expect at 10× and 25× the price. What the set adds is the behaviour of the context layer across that range:
Luna (economy)
Terra (mid-tier)
Sol (frontier)
Quality, without the Graph
64.0
67.4
73.9
Quality, with the Graph
65.1
69.0
75.4
Token Reduction
-47%
-52%
-52%
Context Value Ratio
1.57
1.84
1.92
- The Context Value Ratio improves with model capability, from 1.57 to 1.84 to 1.92. The stronger agent extracts more from the same context, not less. This is an important finding for the Legal Context Graph because it suggests that as foundational models get better and better, their ability to leverage well-designed context foundations only improves. That makes investments in context engineering the type that grows month by month.
- The gains are front-loaded, however. Economy to mid-tier is worth +0.27; mid-tier to frontier only +0.08. A firm does not need the frontier tier to capture most of what the context layer offers — which matters, because the mid tier is where most production work will actually run.
- The mechanism is the same throughout: fewer tokens for the same or better answers — a 47% token reduction on the economy model and 52% on both others. A context layer that lets an agent read less to find the same right answer has savings that grow with how expensive each token is. Note that the token reduction has already reached its ceiling by the mid tier, so Sol’s higher ratio comes from a marginally better quality term rather than from cutting more.
03 – Putting it into practice
What the savings are worth to a firm
A ratio like the Context Value Ratio is helpful in theory but what really matters for law firms and legal departments is how to put it into practice: either to improve their business or to level up their legal service delivery. This section covers exactly that.
Consider a large firm — 2,000 professionals with access to a legal AI assistant — and ask what the Legal Context Graph is worth to them over a year. Three configurations bear on the question, and we measured all three on the frontier model across the same 300 questions:
reasoning effort
cost per question
quality
Search and retrieval only
medium
$0.5037
73.9
With the Legal Context Graph
medium
$0.2682
75.4
With the Graph, effort raised to high
high
$0.4109
79.3
The first two rows are the benchmark result of Section 2, expressed per question: the context layer holds quality and halves the price. The third row is central to what this section is about. Having reduced the cost of a question by nearly half using the Legal Context Graph, a firm need not realise that savings as literal cash. It may instead “reinvest” part of those savings on agent harness configurations that raise accuracy — a larger model, a longer reasoning budget, broader retrieval. For the purposes of this example we take the simplest of these: raising the reasoning effort of the model from medium to high, which permits the same agent to deliberate longer on each question, reasoning at greater length, verifying more and reading further.
What makes that upgrade affordable is the context layer. Raising reasoning effort costs roughly 1.5× more per question whether or not the Graph is present. Applied to the unassisted baseline, it would raise the firm’s spending by about half. Applied on top of a question price that has already been halved, the same upgrade leaves the firm spending 18% less than it does today. The accuracy is therefore purchased with expenditure the firm had already budgeted.
A firm consequently faces a genuine strategic choice rather than a funding request, since neither path requires spending a dollar more than it spends now. The first is to bank the saving while keeping quality held at parity, 73.9 against 75.4. The second is to reinvest it: accuracy rises from 73.9 to 79.3 while spending still falls. The magnitude of each depends on how heavily the firm works the tool, which is the one variable we cannot supply. Across 250 working days and 2,000 professionals:
Questions per person per day
Questions per year
Banked, saved per year
Reinvested, saved per year
Additional correct answers per year
1
500,000
$117,700
$46,400
+27,000
2
1,000,000
$235,500
$92,700
+54,000
4
2,000,000
$470,900
$185,500
+108,000
8
4,000,000
$941,900
$370,900
+216,000
At eight questions per person per day — a realistic figure for a law firm whose professionals genuinely rely on agents as daily work companion — banking the saving returns roughly $940,000 a year, recurring, for a capability the firm has already built. On the other hand, reinvesting it yields 216,000 additional correct answers, with the firm still some $371,000 a year better off than it is today.
The second of those merits particular attention, as it inverts the customary trade-off. Accuracy in AI systems is ordinarily achieved through more spend: a larger model, more reasoning, more retrieval, each of which costs more. Here the ordering runs the other way. The context layer reduces the price of a question first, and the accuracy is funded from the proceeds, so that a firm concludes the year both more accurate and financially further ahead — a combination rarely available.
It matters most, moreover, exactly where a benchmark becomes hard. As a system approaches the ceiling on a question set, the answers that remain outstanding are the expensive ones: the multi-document reconciliations, the questions whose answer appears once in fifteen million characters. Those are precisely the answers that further deliberation buys. A firm that has freed up budget can afford to spend it on that tail; a firm that has not, cannot.
04 – Over the life of a matter
When the matter keeps changing
Everything above is a snapshot: freeze the matter, build the context, ask the questions. That is where most benchmarks stop, but that is not the reality of most legal matters. Documents arrive, get revised, get superseded, and a context layer has to be built and then kept.
The first consequence is that the Context Value Ratio is better understood as a curve than as a single number. Context is built once per matter; the value it returns accrues once per question asked or task performed. A matter that closes quickly, or that a legal team never has cause to interrogate, will not repay the investment however well the context was engineered. As the number of questions rises, the build cost amortized across them falls toward zero and the ratio approaches the answering-only figure reported above. The key question is therefore the break-even agentic volume: the number of questions answered by agents per matter at which the fully-loaded ratio — build cost together with answering cost — reaches 1.0. It is worth isolating because it converts a question a firm cannot answer, is context engineering worth it, into one it can: how heavily do we work a matter?
Three variables set the break-even agentic volume. The first is the cost of building the context layer, driven by the choice of model and by the volume of text processed. We deliberately do not publish our own figure: across our ten matters it varied roughly sixfold on identical settings, because what moves it is the documents rather than any decision of ours, and a single number carried from our corpus to another firm’s would be misleading. This is the one input a firm must measure for itself. The second is the saving per question, which is what the benchmark measures.
The third property warrants closer examination: variations in the techniques used to build the context layer in the first place. Some context engineering approaches, for example, derive their value from a structure that is computed in such a way that every new document added or changed invalidates what was already built and the cost recurs each time a filing lands. Other approaches attach to individual documents, so a new document adds work proportional to itself and leaves everything already built untouched. On day one these context engineering tactics can look identical and cost the same; over the life of a matter receiving documents weekly they diverge sharply. A ratio measured on a frozen corpus flatters the rebuilding design, so a firm should establish which kind it is investing in before comparing headline numbers.
There is a fourth consideration that the two-tier result above brings into focus. The build cost and the answering cost do not have to be paid at the same model tier. Constructing a document-level index is a bounded, mechanical task — read one document, describe what is in it and where — and it is often well within the capability of an economical model. Answering a legal question across a matter is often not. A firm can therefore build its context layer with an economical model and spend the resulting savings on a far more capable model at question time.
That asymmetry has a large effect on break-even agentic volume, because the saving per question scales with the price of the model used by the answering agent while the build cost does not.
EXAMPLE: When the build pays for itself
A 200-document matter. The context layer is built once with an economical model, and questions are answered by an agent using a frontier model — the split described above.
Documents in the matter
200
Cost to build the index for this matter (economical model)
$0.80
Saving on each question answered (frontier model)
$0.2355
Break-even volume, in questions
≈ 4
The build figure is a round illustrative number — for the reason given above, what drives it is the documents.
Because the answering model is costly and the building model is inexpensive, a matter repays its context layer almost immediately — and every question after the fourth is pure return. Ask a hundred questions of that matter across its life and the layer has returned roughly thirty times what it cost to build.
Run the same arithmetic with the economical model answering as well and break-even lands nearer 160 questions — still achievable on a heavily worked matter, but a very different investment case. The gap between those two numbers is the point: it is the pairing of a cheap build with an expensive answer that often makes context engineering pay back fastest.
Note that the current version of the Legal Context Engineering Benchmark does not measure change over time, but we already have started testing expansions to the benchmark that factor in sets of matters that unfold in stages, every document carrying its real-world date. We expect to report on this data set separately in the months ahead. In the meantime, the general point holds: a context layer is not something built once but rather cultivated over time to optimise for what a law firm or legal department values.
05 – The invitation
Why we are publishing the method
What we described in this report is not only the internal benchmark that guides the evolution of NetDocuments’ Legal Context Graph but also the methodology that underpins that benchmark — a methodology that we hope any legal department or law firm can leverage to assess investment in the context layer.
To apply this methodology simply follow the steps outlined.
- Hold constant your model and your harness.
- Write down the types of questions your professionals actually ask, and have someone qualified specify what a complete answer contains.
- Change one context engineering technique or capability at a time.
- Price a correct answer before and after.
- And read some of the answers yourself before believing any of the numbers.
- Where budget permits, run the comparison at more than one model tier — it is the cheapest way to learn whether what you have built is a crutch or an accelerant.
At its heart, that loop is how context engineering can compound day after day, year after year, to build a foundation of intelligence and achieve a new standard for legal service delivery.
© 2026 NetDocuments Software, Inc. All rights reserved. NetDocuments is the registered trademark of NetDocuments Software, Inc. in the United States and other jurisdictions. All other trademarks, service marks, and trade names are the property of their respective owners.
See what context is worth to your firm
We’ll run the LCEB on your matters, with your questions, on the agent you already use.


