Work
What a year of this work looks like
The examples below describe the kinds of problems I have worked on recently in workers' compensation and PEO administration. They are written as patterns rather than as client stories, by design.
Document intelligence with provenance
Operating agreements, amendments, policies, and years of correspondence, ingested so that a question is answered from the governing text rather than from memory. Every answer names the document it came from, and the system knows which version is currently binding and which has been superseded.
An answer without a source is an opinion. A source that was amended last year is a liability.
Structured data kept structured
Payroll registers, plan data, and financial figures arrive as PDFs, spreadsheets, and scans. They are extracted into auditable tables and queried with SQL. A language model never summarizes a number it could look up, because it cannot reliably sum, count, or join.
The failure mode of asking a language model for a total is a confident wrong answer.
Identity resolution across legacy systems
The same company can appear under four names across contracts, commission statements, spreadsheets, and email. A registry gives each company and person one canonical record, with every alias linked to it, so that a question about a client returns everything known about that client.
Most "the system doesn't know" problems are really "the system knows it under a different name".
A regulatory corpus
Statutes, administrative rules, and public meeting minutes for the workers' compensation programs a client operates under, ingested with their authority level recorded. Compliance questions are answered from the source text, with the citation, and a rule outranks a memo.
Compliance answers should point at the regulation, not at someone's recollection of it.
Process mapping from evidence
Before any automation, current-state process records are derived from what the systems actually recorded: who touched what, in what order, how often. The people who do the work validate each record. The records describe; they do not recommend. Automation is decided afterwards, one piece at a time, with a measure for each.
Automating a process nobody has observed automates the parts that were already wrong.
Voice and call assistance
Questions can be asked and answered by phone as well as on screen: speech is transcribed, the answer is retrieved from the same sources with the same citations, and read back. Call assistance, surfacing history and documents during a live call, exists as a working prototype and is described here as exactly that.
The person on the call keeps the conversation. The system keeps the context.
Working with the data where it lives
Scanned-document OCR and any processing of sensitive material runs on GPU hardware I own, inside the client's network boundary. Hosted models are used only where the data allows, and never for protected health information or payroll detail. A survey tool measures how much protected information a corpus contains before any of it is moved. Where records may contain protected health information, HIPAA governs the design from the start: identifiers are detected before ingestion, and that material is never sent to a hosted model.
Owning the hardware is the simplest honest answer to "where does our data go".
How it's built
The layers underneath the work above, in plain terms. Only open technologies are named.
- Documents
- Graph-enhanced retrieval over a vector store and a knowledge graph (LightRAG on Qdrant and Neo4j), every chunk carrying its provenance.
- Structured data
- PostgreSQL and DuckDB, queried with SQL. Language models write queries; they do not do arithmetic.
- Relationships
- An entity registry in PostgreSQL: one canonical record per company and person, aliases resolved by rule and by review.
- Regulation
- Statutes, rules, and public minutes converted to markdown with their authority level recorded, then indexed like documents.
- The agent
- Claude working through the Model Context Protocol, calling whichever store a question needs and citing what it used.
- Processing
- OCR and any handling of sensitive material on GPU hardware I own; hosted models only where the data classification allows.
- HIPAA
- Workers' compensation and benefits records can contain protected health information. A survey tool checks a corpus for the HIPAA Safe Harbor identifiers and health context before anything is moved, and is built so that no document text enters a language model's context during the survey. Anything it flags stays on owned hardware.
- Voice
- Telephony through Asterisk, transcription with Whisper, and the same retrieval behind the spoken answer.
- Verification
- Every public page of this site is tested with Playwright before it deploys, including a check that no client-identifying term appears.
Client work is described here in general terms by design. If you want to talk through a specific problem, the best route is a conversation.