Jev now makes the decisions in BlackOps that I used to hand to an LLM
A decision model that answers with numbers instead of prose. It scores my reply targets, checks voice on every publish, filters my feeds and reads replies to brain tasks.

Every product with AI in it has spots where something needs deciding. Is this post worth replying to. Does this draft sound like me. Is this article worth reading, or is it another funding announcement. For most of BlackOps' life I handed those questions to a chat model, asked it nicely to answer in JSON, and parsed whatever came back.
Over the last couple of weeks I've been pulling those calls out one at a time and giving them to Jev. It now runs in four places inside BlackOps. Two of them were LLM calls I ripped out. The other two are features I built on Jev from the start, because once it was in the codebase there was no reason to reach for a chat model to make a yes/no call. And through a Sortie I can carry Jev into anything else I build, which is how three of my own tools dropped their chat model steps too.
What Jev is
Jev is TypeSafe's System One model, and it doesn't generate text. You send it some state (a post, an email, a draft) and a set of questions, and it sends back typed answers with calibrated probabilities, all answered in parallel in one pass.
Three kinds of question:
- a yes/no condition (TypeSafe calls it a noul), and how likely it is to hold
- a pick from a list, with the full distribution behind the pick
- a score on a ladder of ordered levels
The possible answers are fixed before the call. Jev can't hand back something outside the set I defined, so there's no parsing step, no retry when the JSON comes back malformed, and no page of prompt begging a general model to act like a classifier. I get a number and I draw a line under it.
The same input and the same question give the same answer every time. A chat model doesn't do that, and a threshold on a number that moves between runs isn't much of a threshold.
Interactive explainer
A feature needs a verdict. Before, a chat model wrote an answer and my code parsed it. Now Jev returns typed answers and my code applies the thresholds.
- 1InputThe state (a post, a draft, an email) and the questions. I set the possible answers first.
- 2JevAnswers all the questions in parallel, in one pass. It writes no text.
- 3VerdictTyped answers with probabilities. Same input, same answer.
- 4Feature actsCode compares the numbers to thresholds. Then it queues, blocks, archives or closes.
Before: LLM call
gpt-5-mini, asked for JSON
- Send the post with about 2,000 tokens of calibration prose.
- The model writes text.
- Code pulls the JSON out with a regex.
- Scans run in batches of ten. The same post can land on either side of the threshold on two runs.
Now: Jev
one call, six questions
- Prefilter in code: too old? Already replied in that thread? If yes, stop. No call.
- Ask
in_expertise,has_standing,post_type,avoid,reply_priority,hunt. - Verdict:
avoidlowpost_typequestionreply_priorityhigh - Code: no veto, type is allowed, so queue by priority.
Before: LLM call
gpt-5-mini, asked for JSON
- Send the post with the calibration prose.
- The model writes text. Code parses a score out of it.
- One number holds everything. A topic match can push it over the line.
Now: Jev
same call, same six questions
- The post passes the prefilter.
- Verdict:
huntmatchavoidhigh - Code:
avoidis a veto. Nothing else can override it.
Before: LLM critic
gpt-5-mini, asked for a JSON verdict
- Send the draft and the voice rules.
- The model writes a verdict. Code parses it.
- If the call errored, it recorded a flat 50 with no violations.
- The gate blocked the post with nothing to fix. One tweet got "50/70" three times.
Now: Jev
one call per piece of content
- Ask one voice score on a five-level ladder (15, 40, 60, 80 or 95).
- Ask one probability per style rule: does the text clearly break it?
- Code: a rule at 0.7 or above is a violation. A major break caps the score at 49. The pass mark is 70.
- If Jev times out, the post goes through marked degraded.
Before: topic match
no model, no judgment step
- The item matches a topic I follow.
- It goes straight to the inbox.
- About 70% of captured items got deleted on sight.
Now: Jev
built on Jev from the start
- Score the item on arrival, before any inbox.
- Ask
item_type,usable_angle,novel. - Verdict:
item_typefunding news - Code: that type is on the kill list.
Before: nothing
this feature is new
- It never had an LLM in it.
- Jev was already my default for this kind of question.
Now: Jev
reads the reply before anything closes
- Ask: does the reply resolve the ask? Is it a question back? What does it try to do?
- Verdict: resolves the askclear
- Code: only a clear resolution closes the task. Ambiguous, or a Jev error, leaves it open.
Before: hardcoded
my bug logger
- Severity was a hardcoded default.
- An agent could go into the repo to chase a vague report.
Now: the jev Sortie
any tool fires it by name; the key stays on the server
- Ask: how severe? What kind of record? A duplicate of an open one? Enough detail to act on?
- Verdict: enough detailunder 0.5
- Code: under 0.5 means no agent goes out.
| LLM call | Jev | |
|---|---|---|
| Output | Text that code must parse | Typed answers with probabilities |
| Possible answers | Anything the model writes | Only the set I defined |
| Same input twice | Can change | Same answer |
| Parse and retry | Yes, when the JSON breaks | None |
| Price | Pays for generated text | $0.042 per million input tokens, output free |
| Reply targeting | Batches of ten | $0.000074 per post, six answers in 350ms in one live scan |
Answers show as high or low for illustration. Thresholds are the real ones. Writing (drafts, replies, rewrites) stays with the LLM.
What it costs
TypeSafe charges $0.042 per million input tokens, and output is free. TypeSafe's own claim is 40 to 400 times cheaper and 40 to 200 times faster than frontier models on comparable tasks. Those are their numbers, and they say themselves it's probably the high end, so here are mine.
On reply targeting, after I trimmed the payload, scoring one post costs $0.000074. That works out to about $2.22 a month for someone scoring a thousand posts a day. In one live scan Jev came back in 350ms with all six answers filled in. At that price I stopped rationing decisions. I can put one on every post, every feed item and every publish.
None of this means the LLM is gone from BlackOps. Drafting a reply, revising a post, writing anything at all, that's still a writing model. Jev took over the parts where I was paying a writing model to hand me a verdict.
Reply targeting in the Chrome extension
As I work through X search and explore, the BlackOps Chrome extension flags posts worth replying to, based on my hunts (saved descriptions of the conversations I want to be in). It used to ask gpt-5-mini for prose and regex the JSON back out. So I was paying for text generation to get a score, scans got rationed into batches of ten, and the same post could land on either side of the threshold on two runs. The old prompt carried about 2,000 tokens of calibration prose just to make a general model behave like a classifier.
Now each post goes through a deterministic prefilter first (too old, already replied in that thread), and the survivors get one Jev call with six questions:
in_expertise: does this touch something I have hands-on experience withhas_standing: do I have a first-hand decision, mistake or number to add, beyond agreeingpost_type: question, hot take, announcement, promo, engagement bait or personal newsavoid: is this a pile-on, politically charged, or otherwise a bad place to show upreply_priority: how worth replying it is, on a four-level ladderhunt: which of my hunts it belongs to, or none
The routing is arithmetic in code, never another model's opinion. avoid is a veto no matter what else scores. Promo, engagement bait and personal news get dropped. Everything else queues by priority. Every verdict is saved, so I can tune the thresholds against what I actually replied to, and a skipped post tells me the real reason it was skipped.
Two fixes came out of the first days. Jev was judging my expertise without being told anything about me, so an MCP announcement scored 0.38 on in_expertise for someone who has shipped an MCP server and written three posts about it. The payload now carries what I know, pulled from the names of the brains (sets of markdown notes BlackOps keeps for me) attached to each hunt. Then I measured the payload itself. It was 3,248 tokens per post and the post was 61 of them, because I was sending each hunt brief twice. Cutting the duplicate and capping the reply intent brought it to 1,759 tokens, 46% less.
Brand voice check on every publish
Every time I publish or schedule anything through BlackOps (tweets, threads, blog posts, X articles, LinkedIn, Threads, TikTok) it gets checked against the brand voice I set up for that site. That check used to be a gpt-5-mini critic. When the critic's call errored, it recorded a flat 50 with no violations and blocked the post with nothing to fix. One of my tweets got refused three times in a row with the same bare "50/70".
Jev runs the check now, one call per piece of content:
- one score for overall voice on a five-level ladder, from off-voice to on-voice throughout, mapped to 15, 40, 60, 80 or 95
- one probability per style rule, asking whether the text clearly breaks it
A rule at 0.7 or above counts as a violation. At 0.9 it's a major one and caps the score at 49. The pass mark is 70. When something fails, the refusal names the rules that broke, worst first, so I know what to fix. The verdict also comes back in the publish response with the engine, model, score and threshold, so a real pass never looks the same as a check that got skipped.
It judges each piece as the platform it's going to, so a short reply on X isn't marked down for being short. If Jev times out, the post goes through marked degraded instead of blocking me. If it fails fast, the old LLM critic steps in.
The blog auto-revision loop still uses the LLM critic, because that one feeds written suggestions into a rewrite. That's writing, and writing is still the LLM's job.
Content scoring on every feed item
BlackOps content monitoring pulls in RSS items and discovered content for the topics I care about. Roughly 70% of what it captured got deleted on sight. Most of it was on topic, and useless anyway: funding rounds, launch announcements, press releases with nothing to say. Asking whether something was on topic was the wrong filter.
Now Jev scores every item as it arrives, before it can reach an inbox, with three questions in one call:
item_type: funding news, product launch, benchmark, opinion, tutorial, research, press release or incidentusable_angle: could someone working in my domains say something about this beyond summarizing itnovel: is this new, or another outlet's version of a story already captured this week
Types on a kill list are suppressed. The rest rank by usable angle, with novel catching aggregator duplicates. Nothing gets deleted at ingest. A suppressed item is archived with its scores attached so a bad threshold stays visible, and archived items purge after 30 days. What survives becomes a ranked, capped digest, and if nothing clears the bar the digest says so instead of padding itself out.
The domains usable_angle judges against are set per site at /admin/content-scoring. A site without them doesn't ingest at all, because an empty domain list doesn't score neutral. It scores everything down and looks exactly like a quiet feed.
Reading replies to brain tasks
A brain can now keep track of things it's waiting on, like a photo or an answer only one person has, and email that person a reminder. When they reply, Jev reads the reply before anything closes. It asks whether the reply actually resolves the ask, whether it's a question back, and what the reply is trying to do. Only a clear resolution closes the task. Anything ambiguous, or any Jev error, leaves it open, and I can reopen a task that closed wrong.
This one never had an LLM in it. By the time I built it, Jev was already my default for this kind of question.
A portable Jev I can throw anything at
The four features above are Jev wired into BlackOps itself, and the Sortie is how I carry it into everything else.
BlackOps has Sorties, a saved endpoint that keeps the URL, the headers and an encrypted credential on the server. I put the TypeSafe key in one called jev. Anything that can fire a Sortie by name can use it: a script, a scheduled job, an agent in a chat, Claude or Cursor over MCP. It sends a state and some questions, and since firing a Sortie hands back the endpoint's full response, Jev's typed answers come straight back to whatever asked. None of those callers ever holds the key. It's decrypted on the server at fire time and never ends up in a conversation, a skill file or a repo. Every fire is rate limited and shows up in the Sortie's fire history.
So when something new needs a yes/no, a pick from a list or a score, there's no integration to build. I name the Sortie and write the questions. Three of my own tools moved over that way in one day.
My bug logger turns a description typed into chat into a structured record, then sends an unattended agent into the repo to write the fix. Severity used to be a hardcoded default. Now one Jev call asks how severe it is, what kind of record it really is (which catches the feature request I filed as a bug), whether it duplicates something already open, and whether there's enough detail to act on. Under 0.5 on that last one, it gets filed for me to look at and no agent gets sent off chasing "the tag thing is broken again."
My writing check reads drafts against my voice rules. Jev decides whether a pattern is there, and the writing model finds it and rewrites it. It was good on vocabulary right away. On rhythm, like the same list construction used four times in one draft, it wasn't, so those checks stayed with me.
My bill tracker reads my email for hosting charges, API bills and subscriptions and files each one into a spreadsheet. Working out what counts and what kind of expense it is comes down to a pick from a list, which is what Jev does best. The job already knew how to fire a Sortie, so wiring it up took about as long as making coffee.
Where it stands
It's live. Every BlackOps feature Jev runs in saves each verdict, so I can see what it decided on any post, draft, feed item or reply.
The thresholds are still mostly my guesses: 0.7 for a voice violation, 0.5 for an actionable bug, the floors on the feeds. There was no history to cut them against, which is why every one of these features records its verdicts from day one. They get tuned against what actually happened.
Every place Jev runs also has a way out when it's down. Reply targeting falls back to the old LLM post by post, the voice check lets content through marked degraded, feeds let items through flagged unscored, and brain tasks stay open. I found out why that matters right after launch. Every Jev call from the extension was failing on a missing field and quietly falling back to the LLM, and nothing looked wrong because the fallback worked. It logs an error now when a whole scan fails.
Almost every feature I've built has a step that quietly needs judgment, and I'd handed most of them to a writing model. Jev gives me the same answer every time, and a number I can set a line against instead of a paragraph I have to interpret. I keep finding more places to put it, and with the Sortie, any new place just has to fire it by name.
If you want this running on your own publishing, brains and feeds, BlackOps is at blackopscenter.com/start.
Where in what you've built is there a step that needs judgment, that you either hardcoded years ago or handed to a chat model and stopped looking at? I'd like to hear what you find.
š Want more insights like this?
Subscribe to my newsletter for weekly deep dives into frontend development, AI, and productivity
I wrote this post inside BlackOps, my content operating system for thinking, drafting, and refining ideas ā with AI assistance.
If you want the behind-the-scenes updates and weekly insights, subscribe to the newsletter.
Related Posts

I did not want another writing tool
Every AI writing tool opens with a blank box and asks what you want to write about. Then the model invents something, and it reads like it was written by someone who has never done the work. The gap worth closing is not writing. It is the hour between knowing something and having published it.

How I use a BlackOps brain to track my deploys
I asked BlackOps for a brain called Deployment Log, backfilled 34 deploys into it out of the build logs, and it caught a failed build I had never recorded plus a fix I had already made too narrowly.

Claude's Watermark Has a Shelf Life
Anthropic now watermarks Claude's text output worldwide. The mark works. It will also stop meaning anything, because assistance is on its way to being universal and a signal present on everything distinguishes nothing.