AI medical record review turns a disordered stack of treatment records into a dated chronology with page citations in minutes, and it is dependable for that specific job. It is not dependable for clinical judgment, causation, or deciding what matters, and firms that blur those two things create problems that surface at deposition.
Medical records are the worst part of a plaintiff file. They arrive out of order, from five providers, in different formats, with duplicate pages and a fax cover sheet between every visit. Somebody has to read all of it. That somebody used to be a paralegal with a legal pad, and the quality of your demand letter depended on how alert they were around page 400.
That is the actual problem AI solves here. Not speed. Consistency of attention.
What it does well
Chronology building is the core win. Feed it the records and ask for a table of date, provider, complaint, findings, and treatment, with a page citation on every row. What comes back is the spine of your damages narrative, and it took four minutes instead of six hours.
Gap detection is the second win, and it is underrated. Ask which periods have no treatment and how long each gap runs. Defense counsel is going to build their argument out of those gaps, and you want to see them before the mediation, not during it.
Billing extraction works well too. Pull every charge with date, provider, and amount into a table that totals. This is mechanical work that humans get wrong when tired, and it feeds straight into your demand letter.
Duplicate identification saves real time. The same ER visit often appears in three separate productions. Asking for a deduplicated timeline cuts the volume you personally read.
Where it fails
Handwriting is close to hopeless. Physician notes, intake forms filled in at a front desk, anything scanned at low quality. The model will produce a confident transcription of something it cannot actually read, which is worse than telling you it cannot read it.
Clinical significance is the bigger risk. The system will record that a patient reported prior back pain in 2019 without registering that this line is the entire defense to your causation argument. It reports. It does not litigate.
Contradictions get smoothed. If one provider writes that the patient denied prior injury and another documents a prior injury, a summary tends to present both without flagging the conflict. Ask for contradictions explicitly as a separate question, every time.
And it will occasionally produce a date or dollar amount that appears nowhere in the records. This is the failure that ends careers, and it is why the verification step below is not optional.
The workflow I would run
Run the records through in batches by provider rather than all at once. A single provider’s file keeps the context tight and the citations accurate.
Ask for four separate outputs, not one summary. A dated chronology with page citations. A billing table that totals. A list of treatment gaps over 30 days. A list of internal contradictions and anything that undercuts causation. Separating them forces the model to actually perform each task instead of blending them into narrative mush.
Then verify. Every date and every dollar figure that will appear in a demand letter or a pleading gets checked against the cited page by a human. That is the whole safety system, and it costs a fraction of the time you just saved.
| Task | Trust level | Verification needed |
|---|---|---|
| Chronology of dated visits | High | Spot-check 10 percent against citations |
| Billing totals | High | Check every figure that leaves the building |
| Treatment gap identification | High | Confirm the gap dates |
| Duplicate detection | High | None, errors here are harmless |
| Handwritten notes | Low | Read them yourself |
| Causation and significance | None | This is legal work, do it yourself |
The confidentiality question you have to answer first
Protected health information belongs to your client, and running it through the wrong tool is a problem regardless of how good the output is.
Consumer tiers of chat products are the wrong home for it. What you want is an enterprise agreement with zero data retention and a business associate agreement, which the major providers do offer. Get the specific terms of your specific plan in writing. “I read that they do not train on business data” is not the same as having the agreement.
Then check your state bar. Guidance on AI and client confidentiality has moved quickly and varies, and the analysis in AI ethics for lawyers walks through what the opinions issued so far actually require. The recurring theme is competence and supervision, not prohibition.
What this is worth
At meaningful volume, record review is one of the largest labor costs in a plaintiff practice, and it is work that does not improve with seniority. Nobody gets better at reading page 600.
The firms getting real value are not the ones that fired their paralegals. They are the ones whose paralegals stopped transcribing and started verifying, which is a better use of a trained person and produces a better file.
Do this today
Take one closed file where you already know the answer. Run the records through and ask for the chronology with page citations. Compare it against what your team produced originally.
You will find two things: places where the AI caught something your team missed, and at least one entry that is wrong. Both of those are the point. The first tells you it is worth using. The second tells you exactly how much verification your process needs.