LXLegalCase 02

Every document in a case, searchable and cited — across thousands of scanned pages.

The problem

A law firm runs cases with dozens of documents each — many of them scanned PDFs. Lawyers had no way to ask a question and get an answer grounded in the actual filings.

Our approach

Documents are uploaded per case; OCR extracts the text from scans and PDFs; the content is chunked and embedded into a vector database. Lawyers then query the case in plain language and get answers that cite the exact document and page — nothing invented.

User journey
Upload case docs
OCR extract
Chunk
Embed
Vector DB
Cited Q&A
System architecture
Azure · per-tenant isolated
Ingress
Entra ID auth
Blob uploadSAS · per-case
upload event
Event
Event Grid
Function: dispatch
fan-out
OCR
Azure Document Intelligencescanned PDF → text + layout
text extracted
Index
Chunksemantic + overlap
Embedprivate endpoint
Azure AI SearchHNSW + hybrid
RAG agent
Retriever tool
Reranker
Answer agentcites or refuses
grounded, with citations
Output
Cited Q&A UI
Per-case audit log
Scales viabatch OCR workers ×Nvector index · HNSWhybrid (keyword + vector)per-case isolation
OCRChunkingEmbeddingsVector DBRAGCitations
Thousands of pages indexedCited answers0 hallucinated clauses

Representative build — real project type and architecture, anonymised client, illustrative figures.

Want this for your stack?

I'll map your problem to an architecture like this — and prove it with an MVP before you pay.

Start a build →

Next: InspectionCracks, spalling, and paint failure