Aloi’s judgment graph for law firms: verification checklist, RFP and pilot plan

Turn firm memory into decisions, but don’t buy the pitch without proof

Picture a senior partner’s instincts, a mental library of “how we do this” turned into recommendations a junior lawyer can pull up on demand. That’s Aloi’s claim: capture firmwide legal judgment from large document stores and surface context-sensitive recommendations for drafting, negotiation and client preferences.

This write-up relies mainly on Johan Häger’s interview on Artificial Lawyer TV and Aloi’s public framing. Treat the vendor’s statements as claims to be verified: require demos with provenance, measurable benchmarks, and security attestations before you put any machine-generated “judgment” into production.

What Aloi says it does, and the verification questions you should ask

According to CEO Johan Häger on Artificial Lawyer TV, Aloi positions itself as a data-processing layer sitting above a firm’s DMS and below front-end workflow apps. Aloi says it ingests large, unstructured corpora, applies model specifications and hundreds of metadata fields, tracks version lineage, and learns deal context via what they call a “judgment graph” to surface recommendations.

  • Scale and ingestion: Aloi says it has worked with examples such as “a law firm with a hundred million documents” and that technical onboarding can take “a couple of hours” while ingesting “a number of million documents over a few days.”

    Ask: provide the ingestion benchmark: raw vs deduplicated document counts, documents/day, hardware/network profile, and a sample SLA. Which steps take “a couple of hours” and which take days?
  • Structure and engineering effort: Häger said Aloi invested significant legal-engineering work to teach the system to recognize document structures, rules, dependencies and relationships.

    Ask: produce a model-spec sample and a short description of the metadata schema (show examples of the “hundreds” of fields).
  • Version lineage and negotiation data: Aloi highlights version handling (e.g., an SPA with “20, 30 different versions”) as a signal for learning how clauses evolve.

    Ask: demonstrate lineage linking on a live matter: show the draft timeline, edits, and how those edits influenced the recommendation.
  • Marginal returns: Aloi says value rises steeply early (10→100 examples) and flattens later (1000→10, 000).

    Ask: request the vendor’s marginal-value curve by clause type and the inflection point where precision gains flatten.
  • Customer scale: Aloi says small firms (around 30-40 lawyers) have reported good results and that they’ve done bespoke integrations (they referenced a team sent to set up a middleware/MCP connection for a client).

    Ask: supply customer references or anonymized case studies that quantify time-saved and provide integration examples (which DMS, what connector work was done).

What a “judgment graph” really means, and what to demand

“Judgment graph” is a useful label, but don’t treat it as mystical. Ask the vendor to specify whether it is:

  • a labeled knowledge graph linking clauses, versions, outcomes, parties and negotiation metadata;
  • a vector index augmented with rich metadata and version lineage; or
  • a rules/heuristics layer mapping precedent → recommended decision.

Demand a schema example: show nodes, edge types, and how confidence is computed. If the vendor cannot show a concrete schema and examples, treat “judgment graph” as marketing shorthand, not a verified capability.

Why firms should care

Moving from “find precedent” to “recommend a firm-specific approach” addresses three real needs: consistent client advice, faster ramp for juniors, and less reliance on a handful of senior partners for repeat transactional work. If provenance, security and accuracy are sound, the judgment layer can increase client stickiness and operational efficiency.

Governance and security, exact asks every buyer should make

Legal work demands stringent controls. Here are specific acceptance criteria to include in your RFP or vendor conversations:

  • SOC 2 Type II report for the most recent 12-month period, plus an executive summary of findings and remediation items.
  • External penetration test results within the last 12 months and a remediation timeline for any critical findings.
  • Data Processing Agreement (DPA) aligned with GDPR, including data retention, deletion, and incident notification timelines.
  • Customer key management (BYOK/HSM) or equivalent for sensitive client data; explain whether encryption keys are customer-controlled.
  • Training‑data policy that answers: do you train models on customer data? Are fine‑tunes isolated per customer? Can customer data be removed (“unlearned”) on request?
  • Logical/physical separation options (dedicated tenancy or VPC) for high-sensitivity clients.
  • Conflict and privilege controls: documented filters to prevent cross-client leakage and privileged content exposure.
  • Provenance and explainability: every recommendation must be accompanied by matching source clauses/matters, version lineage, and a confidence metric showing why those sources were relevant.

Evaluation / RFP checklist you can paste into procurement

  1. Architecture whitepaper: Provide an architecture diagram that details model types (open LLM, proprietary, hybrid), storage/encryption model, metadata schema sample, and the judgment-graph representation. Acceptance: whitepaper + Q&A session with technical leads.
  2. Provenance demo (live): For a real or anonymized matter, show three recommendations, each with the exact source clauses, version timeline, negotiation notes, and a confidence score. Acceptance: at least 2 of 3 recommendations must link to legally relevant sources as judged by your counsel.
  3. Accuracy benchmarks: Provide precision/recall for clause detection and recommendation relevance on a labeled hold-out dataset. Include dataset size, labeling methodology, and inter-annotator agreement. Acceptance: vendor sets baselines; for production use expect high precision thresholds (e.g., ≥90% for clause detection where false positives are costly) or a documented mitigation plan.
  4. Ingestion and throughput metrics: Show documents-per-day benchmarks, deduplication rates, and typical CPU/storage footprint for X million documents. Acceptance: vendor provides an ingestion plan with SLA and resource estimates for your dataset size.
  5. Security/compliance artifacts: SOC 2 Type II, recent pentest report, ISO 27001 if available, DPA template. Acceptance: review by your security/compliance team; unresolved critical findings are a blocker.
  6. Customer evidence: Two references or anonymized case studies that quantify time-saved, junior ramp-up, or retention impact, with contactable references or signed attestations. Acceptance: at least one reference in a similar practice area and firm size.

Pilot template, a practical path to proof

Run a structured pilot before any roll-out. A focused pilot reduces risk and creates clear go/no-go criteria.

  • Duration: 8-12 weeks.
  • Scope: one high-volume transactional practice (e.g., corporate M&A or facilities agreements), covering ~200 transactions or ~100k documents, depending on firm size.
  • Participants: 3-5 senior lawyers, 6-10 junior lawyers, and 1 technical lead from IT/compliance.
  • KPIs: percentage reduction in drafting time (target ≥20%), percentage of recommendations accepted without edits, reduction in review cycles, and junior ramp-up time to a defined competency level.
  • Governance guardrails: human-in-the-loop required for all client-facing outputs, rollback triggers for confidence < X%, and weekly review of provenance logs.
  • Exit criteria: inability to demonstrate provenance for ≥75% of recommendations, unresolved critical security findings, or KPI shortfalls vs agreed thresholds.

Where this will help first, and where to be skeptical

Most promising: repeatable transactional flows with many near-duplicate clauses, NDAs, standard SPAs, facility agreements, and recurring procurement contracts. These areas produce the volume and repeatability needed to learn stable decision patterns and measure time savings.

Less promising: one-off regulatory strategy, bespoke litigation theory, or novel cross-border arbitrations. Low repeatability means fewer reliable precedents and lower coverage. For such work, expect significantly lower match rates and require manual escalation paths.

Key questions, with short, honest answers

  • Can a system actually capture legal judgment?

    According to Aloi’s CEO, they combine model specs, extensive metadata, version lineage and a judgment graph to generate recommendations; treat this as a vendor claim until you see provenance, accuracy metrics and a live demo linking recommendations to source matters.

  • How much data do you need for value?

    Aloi says small firms (~30-40 lawyers) can see value and that marginal gains rise quickly at low counts (10→100 examples) then taper; ask the vendor for their marginal-value curve by clause type and for a sample coverage report on your corpus.

  • Will this replace junior lawyers or commoditize work?

    Häger predicts pressure on low-value drafting pricing. That’s a plausible market dynamic; firms should pilot, measure savings, and redesign staffing and pricing models deliberately rather than reactively.

  • What governance must a firm require?

    Insist on data isolation, an explicit training-data policy, provenance for recommendations, conflict filters, SOC 2 Type II and a recent external pentest. Those controls are non-negotiable for client confidentiality and ethical practice.

Next moves for legal leaders

  • Ask Aloi (or any judgment-layer vendor) for a provenance demo you can evaluate with inside counsel present. Do not accept abstract descriptions without live examples.
  • Demand SOC 2 Type II, a recent external pentest, and a DPA with clear deletion and incident notification clauses.
  • Run the pilot template above with agreed KPIs and rollback triggers; make a commercial decision based on measured outcomes, not aspiration.
  • Prepare pricing and staffing experiments in parallel, if low-value drafting becomes automated, plan how to redeploy experienced lawyers to higher-value advisory work and reflect that in billing models.

Turning institutional legal judgment into a usable layer is a sensible engineering challenge. Aloi’s stack, model specs, metadata, lineage and a judgment graph, maps to that ambition. The decisive question for buyers is never whether a vendor promises judgment, but whether the vendor can prove it with auditable provenance, rigorous metrics, and ironclad governance. Ask for the receipts; make decisions on the demos and data, not the marketing.