Hi, My Name Is Jev: Decision Models and Clinical Quality
Clinical quality has always had a measurement problem. Not only what do we measure, but how?
We've had chart review, inter-rater reliability, and scoring rubrics for a long time. The catch has always been that we need humans to do the measuring, which is fine in small batches - but unwieldy at scale. Decision models like Jev give us a new way to measure at scale on terms that we set - not by guessing on an LLM's interpretation, and not by trying to scale up human beings. Done right, that can shorten the feedback loop on treatment quality from quarters down to sessions - and give us new ways to understand what quality therapy looks like.
What is Jev?
Jev is a decision model, not a generative one. You don't ask it for an answer - you ask it a narrow question, and it gives you back a probability: how likely something is to be true or false, how likely it belongs to a particular category, or how it should rank. Yes, this was in there. No, it wasn't. Here's where it most likely belongs.

Daniel Miessler has a much better explainer of how decision models work than I'll attempt here, so I'll point you there for the specifics.
What makes decision models so exciting right now is that, like LLMs and agents, they're able to parse a large amount of text and are accessible via API. This means they're easy to call and integrate into software and other workflows. But unlike LLMs, they cost much, much less - and can respond in milliseconds. That means we can take the same data we'd normally send to an LLM, break it into smaller questions, and get faster answers - each with a measurable probability attached.
Why not just ask an LLM?
"Is this good therapy?" is too open-ended a question for an LLM. And if you try to parse it finer, you hit a different wall - an LLM can come back with a different answer every time you ask. Most of the common models are generative, not deterministic - that's what they're designed for. (Meaning, it's difficult to get them to respond consistently, or to even understand where their answers are coming from.) They're also expensive on a per-token basis: around $2-4 per million input tokens and $10-20 per million output tokens, which adds up quickly at scale.
For clinical quality, their unpredictability becomes an auditability problem. An LLM might be able to tell you whether something exists in a transcript, but you can't really check whether it's right without going into the transcript yourself. And a confident answer from an LLM can touch every one of your criteria without ever answering the specific, variable-based questions a quality measure needs. That makes aggregate analysis much harder.
You can't audit a paragraph. But you can audit a probability and an agreement rate.
How certain the model in its answer needs to be is a threshold we define. We already have solid measures for interrader reliability with humans. So how can we extend that to working with AI? This is where a decision model comes in, as it gets you a few things an LLM output doesn't:
- You can ask much more fine-grained questions.
- You set the threshold for what counts as a yes, instead of taking whatever threshold the model decides on.
- You can check it against human raters, and route what's uncertain into a human-in-the-loop process you designed.
That threshold piece matters more than it sounds. The decision about what "good enough" looks like becomes explicit, and it sits inside your organization's control. The same transcript can generate very different probabilities depending on the kinds of questions you ask; whether a risk assessment was conducted or whether the therapist told the patient what to expect at their next session. The answer to both questions matters. But the risk assessment is the one you need to happen every time, with a high probability that it did - it can also factor into thigns like HEDIS or Stars measures for providers. The next-session detail can live with a looser bar.
Those same thresholds also let you put actual numbers on human-in-the-loop. As much as we talk about wanting to include people in the process, we know it can't happen every time without risking automation fatigue. So how do you decide, with statistical regularity and rigor, what likely needs a human and what doesn't? You need something That you define, that's also auditable and defensible for when something goes wrong.
Why do we wait 30, 60, and 90 days?
Today, clinical quality is typically measured at 30, 60, and 90 days. That's because that 30/60/90 cadence is doing three jobs at once.
First, it's based on clinical timing. A patient sees a therapist every week or every two weeks, and that window lines up with what we know about the course of treatment, and how long it takes to see a minimal clinically important difference on something like the PHQ-9 or GAD-7.
It's also an artifact of in-person care and paper record review. Giving a patient a PHQ-9 every session is probably overkill, so it usually goes out around once a month to track the trajectory of treatment. Across 30/60/90 that gets you somewhere between three and six measurement points, depending on how often it's given.
Lastly, it's about statistical power. You need enough data points to spot trends within a patient and across groups of patients.
The result is that quality data shows up as a lagging indicator, not a leading one. It just takes that long to gather enough to analyze.
Meanwhile, a lot happens inside any 30 days. That's two to four sessions where the therapist and patient are talking about and doing a lot of things. So there are a lot of data points in there you can't capture without asking them to do the extra work of all that data entry. That's also part of the problem with the current PHQ-9/GAD-7 model - it's a lot of questions to ask a patient every time just to track where they're headed.
What if you asked a lot of narrow questions of the data that's already being collected instead? How quickly a patient signs up for treatment? Whether they show up? What's said in the room, by both the patient and the clinician? A model like Jev can make many narrow decisions from that data and try to get upstream of the 30/60/90 measurement. More questions also means more statistical power, so you can see trends in aggregate sooner.
Right now we look at things like outcomes, time to treatment, or net promoter score as ancillary measures of what good therapy would look like. But what if we could measure treatment adherence or therapeutic relationship without asking the patient or the provider to do anything extra? Asking more targeted questions about the things that would happen in therapy regardless, and putting measurements against them. But I can't say yet whether those measurements actually gain us anything. They'd need to be matched against the lagging data to see which process judgments are worth keeping.
What happens when you can zoom in on the loop?
So let's take a more practical look at what this could look like. Clinical ops leaders want to know a lot of things. Are my clinicians doing evidence-based treatment? Are they catching things like risk assessments and handling them appropriately? Are they doing good therapy? And in a digital, on-demand business, what does good therapy even look like, and how do we measure it?
Again, answering those questions today means waiting 30, 60, or 90 days just to submit the ticket to the data team. And if the something isn't quite right, you get back in the queue for the next three to six months to have it updated. So instead of being able to do our course-correcting monthly, we do it in quarters or years.
What a decision model gives us is the same loop at different speeds and levels of zoom.
The encounter.
If we can look at quality inside the session, we can adjust while still inside the session. Catching whether a therapist asked about medical diagnoses, current medications, or ran a suicide assessment - and prompting them to do it before they close - lets you spot the problem and solve it in the same beat. You can't do that quickly with an LLM. With a model that responds in milliseconds, you can course-correct as things happen, at the individual level, instead of waiting six to twelve months to look at it in aggregate or saving it for an annual in-service.
The episode.
Across an episode of care, maybe we can find triggers that correlate with outcomes or therapist quality. Right now our main tools are the GAD-7 and PHQ-9, which are paper-based, cumbersome, and irritating for people to answer in an on-demand world. What if we could find measures that don't require asking anything?
The organization.
This is where decision models can pair nicely with something like Hex, which product teams already use to pull data faster without getting into the data science queue. If adding a new question costs next to nothing, and it can be instantly rerun on existing data plus everything going forward, how does that change the way we tune our systems and our questions? That kind of question-asking used to belong to software and product folks. Clinical ops could have it too.
That flexibility also means clinical quality leaders have to get more disciplined, so we don't p-hack our way into findings that look significant and don't matter at all. With new tools come new responsibilities, and one of those is getting tighter on our data processes as we open them up to ask more questions.
Who actually benefits?
There's always been a tension at the center of clinical quality. Over-index on standardization and therapy starts to feel rote or alien - not the humanistic work most therapists got into this for. Frankly, if you standardize therapy too much, what the heck do you need therapists for except to deliver it?
Go the other direction, with every therapist doing whatever they feel like, and it's hard for payers and employers to quantify what they're paying for. The danger there is they don't pay, and we lose access for the people who need it most. And I think that's what we've seen in the current paradigm.
Networks manage that tension today with regular audits, most of them manual. And manual is slow. When I was at Teladoc, we estimated that manual chart review of every provider on the platform, one by one, would take two full years. By then we'd have cycled through too many new providers for it to make a difference.
AI has already started to change that, with LLMs analyzing notes for audits and chart review at scale. I think decision models are the fast follow - a faster, cheaper, more efficient way to do the same thing: automated quality measurement at scale.
What is in this for clinicians?
Faster feedback than 30/60/90 on what they can do to put a patient on the right track and keep them engaged, and most importantly without doing anything extra. Clinicians could get something like a post-session report - a personal dashboard of how they're doing across sessions - or even a reminder in-session of something they forgot to ask that might be important. And we can all get quality asks that feel less arbitrary. Right now, quality decisions can seem arbitrary coming from the company, or even punitive when they arrive as demands from payers and employers. If those decisions are grounded in data that shows how they affect outcomes, it becomes something we can all get behind.
What is in this for the organization?
An understanding of what actually leads to better outcomes, which is what ties back into value-based contracting. It also lets you spot, for lack of a better term, "super shrinks" - who the top performers are and, more importantly, why. Understanding what your best therapists do differently, without annoying them with continual surveys and interviews, can improve quality across the rest of the network and shape hiring and retention. It gives clinical ops a tighter focus on what to standardize and where therapists can be more idiosyncratic.
What is in this for payers?
A more data-driven quality process that measures performance faster than 30/60/90 and fits better with how they already think about episodes of care. It also gives them more options when defining value-based contract terms - more than the standard engagement, NPS, and outcome measures, and based on things that actually tie to outcomes, not what we think ties to outcomes.
What has to exist first?
For any of this to become real, you need a decision model that is clinically validated and HIPAA compliant. Jev, the most popular one, isn't there yet - it has no BAA available.
Validation also means actually training the model against clinician raters and building it from your own data. Start with what successful outcomes look like. Find the patients who match that. Measure every data point across their trajectory, and then see how well the model holds up against the rest of your population.
But until we get a HIPAA compliant and clinically validated version, all of this remains a "what if" exercise...
More to explore
What Comes After the Note
The clinical note asks one document to serve memory, treatment, and billing — all from the therapist's recall. When sessions start producing transcripts, biometrics, and semantic signals, what replaces the note, and who takes clinical responsibility for it?
How do you go from zero to 10x with AI as a therapist?
A 16-week, competency-based curriculum to help mental health clinicians go from zero to 10x with AI, covering foundations, clinical workflows, prompt engineering, assessment, ethics, and governance.
AI and Psychology: What are the big questions for 2025?
Key takeaways from the APA Mobile Health Tech Advisory Committee meeting on AI in mental health, exploring critical questions about the role of psychologists, AI tools, ethics, equity, and explainability.
Enjoyed this? Get new essays when they're published.