A contract assistant finds three notice periods: 60 days, 90 days and 30 days. All three numbers appear in the supplied documents. Which one should it return?
Useful contract AI connects a clause to the right service, date and document status. This walkthrough gives teams a practical way to check those connections: 12 synthetic contract families, with two questions for each family. The examples are deliberately short enough to inspect by hand. The aim is to help teams evaluate and test contract AI with clear, traceable evidence.
This walkthrough uses invented contract rules and includes a preliminary browser pilot. The examples are reading-comprehension exercises, not legal advice or templates for real agreements. The answer key remains provisional; the pilot’s methods and limits appear alongside its findings.
Consider this fictional document family:
Scenario 1: Renewal notice. Three illustrative contract excerpts and paired questions.
Scenario 1: Renewal notice. Illustrative contract excerpts. Open full-size image.
Ask: “For the Analytics renewal dated 1 August 2026, how many calendar days of advance notice does the supplied record require?”
The expected answer is 90 days. The original 60-day term has been replaced for that renewal. The later 30-day proposal does not take effect under the fixture’s explicit rules.
Change only the renewal date to 15 March. The expected answer becomes 60 days because the amendment’s stated effective date has not arrived.
A system that always chooses the latest document would miss the August question. A system that always selects the signed amendment would miss March. Paired questions help expose a selection rule that happens to produce one correct answer.
The original 60-day clause is genuine within this fictional record. Citing it for an August renewal would still leave the answer unsupported unless the response also resolves the amendment.
Check three things separately:
The third check requires the change rule, the amendment’s effective date and replacement language, and the proposal’s status. A response may properly cite an older clause to explain why it no longer governs the question. Evaluate how the citation is used.
The teaching pack contains 36 source documents and 24 questions. These are the scenarios and provisional expectations:
Renewal notice: a signed amendment and later unsigned proposal. August renewal: 90 days. March renewal: 60 days.
Service-credit ceiling: successive amendments replace the same ceiling. July: USD 1,500. April: USD 2,000.
Scenario 2: Service-credit ceiling. Three illustrative contract excerpts and paired questions.
Scenario 2: Service-credit ceiling. Illustrative contract excerpts. Open full-size image.
Different provisions: separate amendments change payment and incident notification. May payment: 45 days. May incident notice: 24 hours.
Future fee: an amendment is signed before its effective date. June: USD 1,000. July: USD 1,200.
Scenario 4: Future fee. Three illustrative contract excerpts and paired questions.
Scenario 4: Future fee. Illustrative contract excerpts. Open full-size image.
Midnight boundary: a response-time amendment starts at an exact timestamp. Before the boundary: 8 hours. At the boundary: 4 hours.
Scenario 5: Midnight boundary. Three illustrative contract excerpts and paired questions.
Scenario 5: Midnight boundary. Illustrative contract excerpts. Open full-size image.
Storage allowance: an amendment is uploaded before its future start date. September: 100 GB. October: 200 GB.
Scenario 6: Storage allowance. Three illustrative contract excerpts and paired questions.
Scenario 6: Storage allowance. Illustrative contract excerpts. Open full-size image.
Order exception: an extended payment period belongs to ALPHA only. ALPHA: 60 days. BETA: 30 days.
Scenario 7: Order exception. Three illustrative contract excerpts and paired questions.
Scenario 7: Order exception. Illustrative contract excerpts. Open full-size image.
Unnamed service: two tiers have different availability targets. Unspecified service: ask which one. Standard: 99.5%.
Scenario 8: Unnamed service. Three illustrative contract excerpts and paired questions.
Scenario 8: Unnamed service. Illustrative contract excerpts. Open full-size image.
Separate schedules: a reporting change affects NORTH only. NORTH: weekly. SOUTH: monthly.
Scenario 9: Separate schedules. Three illustrative contract excerpts and paired questions.
Scenario 9: Separate schedules. Illustrative contract excerpts. Open full-size image.
Unsigned draft: internal approval does not satisfy the fixture’s signature rule. Operative fee: USD 5,000. Proposed fee: USD 4,000.
Scenario 10: Unsigned draft. Three illustrative contract excerpts and paired questions.
Scenario 10: Unsigned draft. Illustrative contract excerpts. Open full-size image.
One signature: only one party signed a proposed change. Operative delivery: 5 days. Proposed delivery: 2 days.
Scenario 11: One signature. Three illustrative contract excerpts and paired questions.
Scenario 11: One signature. Illustrative contract excerpts. Open full-size image.
Missing signature page: execution status is explicitly unknown. Current retention: request execution evidence. Original retention: 30 days.
This list summarizes the cases. Supply each case’s source documents during a test; do not give the model these expected answers. Dates, units, boundaries and document rules must travel with the questions.
The third case starts with payment at 30 days and incident notification at 48 hours. A February amendment changes payment to 45 days. An April amendment changes incident notification to 24 hours. Each amendment explicitly leaves other provisions unchanged.
For the May questions, the expected terms are 45 days and 24 hours. April’s document does not erase February’s payment change just because it is newer. The unit of reasoning is the provision and its scope.
Scenario 3: Different provisions. Three illustrative contract excerpts and paired questions.
Scenario 3: Different provisions. Illustrative contract excerpts. Open full-size image.
Sirion’s explanation of execution dates and effective dates provides background for another distinction in the pack: signing a change and starting an obligation can be different events. The fictional text here is original and does not test Sirion’s software.
An unsigned proposal and an incomplete scan are different inputs.
In the unsigned-draft case, the record explicitly says the proposal was not signed. In the retention case, the signature page is missing and execution status is unknown. The record does not establish that an executed copy never existed.
Under that second fixture, neither 30 days nor 90 days is established as the single operative answer. Request the executed copy or verified execution status and explain the conditional alternatives. The filename retention-amendment-final.pdf does not resolve the uncertainty.
Pair that question with “What retention period did the original agreement specify?” The expected answer is 30 days. This catches an assistant that refuses every question whenever it sees an incomplete document.
Scenario 12: Missing signature page. Three illustrative contract excerpts and paired questions.
Scenario 12: Missing signature page. Illustrative contract excerpts. Open full-size image.
To illustrate the evaluation method, we used the general-purpose ChatGPT and Gemini browser applications. No CLM platform or Sirion product was tested. On 20 September 2026, we submitted the 24 questions once each, producing 48 completed responses. Every question received its full three-document family in a fresh temporary chat, with the same instructions and no answer key.
AI-assisted review against the provisional key found the expected substantive conclusion in all 24 responses from each app. Both asked which service was intended in the unnamed-service case. Neither invented a cited clause ID. These observations describe this small exercise, not a general accuracy rate.
The missing-signature case provided a useful implementation check. Both explanations said the supplied record could not establish whether retention was 30 or 90 days and requested execution evidence. But ChatGPT returned "decision": "answer" alongside a null value. Gemini returned "decision": "insufficient_evidence".
Observed ChatGPT response to F12.Q01: the explanation acknowledges missing evidence, but the decision field says answer.
Actual browser response, 20 September 2026. Open full-size ChatGPT screenshot.
Observed Gemini response to F12.Q01: the decision field says insufficient_evidence.
Actual browser response, 20 September 2026. Long JSON lines extend beyond the visible code panel. Open full-size Gemini screenshot.
Both explanations recognized the missing evidence. The additional check concerns how that uncertainty is passed to the next step: ChatGPT’s decision label did not match its explanation. For a contract workflow, validate that the decision, value, explanation and requested evidence agree before using the response to trigger an action. This observation concerns one browser response; it does not demonstrate a defect in a CLM product.
Methods and limits. This was a single-run, full-context browser pilot, with no retrieval comparison. ChatGPT used High effort; an initial GPT-5.6 Sol selection reverted to Latest in fresh chats, so a fixed model version was not established. Gemini displayed Pro, with 3.1 Pro shown in its mode menu at the start and end. The same AI agent helped create the fixtures and assess the outputs; neither the key nor scoring has independent specialist sign-off. Exact prompts and outputs were retained, with screenshots for this example. The findings do not evaluate Sirion or establish production reliability.
First give the model all three documents in a family and one question. Keep the answer key out of its input. Use a fresh context for each question and instructions such as:
Answer using only the supplied fictional record. Apply its explicit rules. Identify the relevant service and date. Cite supporting clause IDs. If material evidence is unknown, identify what is missing rather than assuming it. Do not treat document order, upload dates or filenames as precedence rules.
Then introduce retrieval. Compare ordinary text retrieval with retrieval that also uses dates, scope and document relationships extracted from the same corpus.
Hold the generator model, questions, prompt and access to source documents constant. Record the chunks actually passed to the model. If humans label the metadata, disclose that extra supervision. Do not give one condition the correct answer through an “applicable” flag.
If full-context answering succeeds but retrieval-based answering fails, inspect the retrieved evidence before changing the prompt. If both fail, compare the response with the source rules before blaming retrieval.
Record the question, exact model version, settings, corpus version, supplied context, raw output and execution errors. Assess the decision and value, evidence coverage, nonexistent citation IDs, and appropriate requests for clarification. A correct number with an incorrect scope explanation should not pass a string-matching check.
Keep operational errors in the report. Show correctness among completed calls and correct answers divided by all planned calls. Preserve reviewer disagreements instead of silently changing the answer key after seeing outputs.
Repeated calls on 12 related families are not hundreds of independent examples. Even cases reserved for later evaluation were visible to the pack’s creator; they are not a blind external holdout.
FiscalQA Pro studies retrieval of date-applicable versions in French tax law. TIDE examines temporally evolving documents and amendment timelines in customs instruments. Legal version resolution is an existing research problem.
This pack proposes a compact commercial-contract exercise combining dates, service scope and execution evidence. Its short, explicit documents make the reasoning inspectable and easier than many production settings.
It does not cover OCR errors, long negotiated contracts, multilingual documents, access permissions or unresolved legal ambiguity. Passing it would not establish legal reliability. Failing it would identify an example to investigate, not a population-wide failure rate.
The practical takeaway is to evaluate the whole answer: the clause it relies on, why that clause applies, and how uncertainty is represented. The browser pilot shows how a small, inspectable exercise can inform a workflow check even when the substantive conclusions are appropriate.
Start with the renewal example. Ask both date-specific questions, retain the outputs, and check whether the response identifies the clause, service, relevant date and evidence. That is a concrete first test a developer can perform before extending the suite.