\n\n

What happened when I pointed Jev, SALI, and CLOC’s Core 12 at every post we’ve ever written

I read Nate Jones’s post on Jev, TypeSafe’s new “System One” model, and he described it as an LLM that can only talk in multiple choice. It reads what you give it, you hand it the possible answers, and it picks one and tells you how sure it is. It can’t write you a sentence. My first reaction was, “What good is that?” My second reaction, about ten minutes later, was, “Wait… that’s what a cataloger does all day.”

So I did what any recovering law librarian would do with a new classification toy. I grabbed the biggest pile of legal-industry text I have the rights to, which is every post on 3 Geeks and a Law Blog going back to July 2008 (2,269 posts, about 3.9 million words, including 84 Geek in Review episodes with full transcripts), and I tried to catalog all of it twice. First against the SALI Legal Matter Specification Standard, then against CLOC’s Core 12 legal operations functions.

Here’s a link to the spreadsheet with the SALI and CLOC tags.

Let me walk you through how it went, warts and all.

The Setup

The division of labor is the whole trick, so let me explain it before the numbers.

Jev did the first pass. It’s cheap (about four cents per million input tokens, with output free) and fast (about a quarter of a second per call in this run). I called it through OpenRouter, since TypeSafe paused new signups the day I tried to get a key. Every call asks Jev a batch of yes/no questions about one post (“Is Knowledge Management a substantial topic of this post?”), and Jev returns a probability for each one.

Claude did the second pass, and it did it inside my regular Cowork session, which meant no per-call API charges. Any tag where Jev was unsure went into a review queue. Claude read the post and kept a tag only if it could quote a sentence, word for word, that supported it. A script then went back and checked every one of those quotes against the actual post text.

Code did everything else. It enforced the rules that a child tag can’t survive without its parent, capped each branch at five tags per post, and merged everything into one spreadsheet.

Before touching the blog, I also ran a warm-up on the podcast RSS feed (370 episodes). That’s where most of the early mistakes happened, and I’ll get to those.

Run 1: SALI

SALI’s standard is a family tree of 18,324 tags, published as an open ontology file on GitHub. I used four branches: Area of Law, Service, Document / Artifact, and Legal Use Cases (about 1,200 tags in all). Jev can only answer the questions it’s handed, so the script walks the tree like a game of 20 Questions. It asks about the 31 top-level Areas of Law in one call, then goes one level down only where Jev said yes, then one more level.

The calibration step. Before running everything, we pulled a 100-post sample spread across the years, including 25 transcript episodes, and had Claude review every tag Jev scored 0.30 or higher. That’s 1,610 judgments, made without Claude seeing Jev’s score. This is the part that saved the project, because it showed Jev’s confidence scores meant something different than I assumed:

Jev’s score How often it was right
0.30 to 0.49 9%
0.50 to 0.69 20%
0.70 to 0.89 41%
0.90 and up 84%

On the podcast warm-up, I’d been letting Jev auto-keep anything at 0.70 or higher. On blog posts, that rule would have been wrong about half the time. The Document / Artifact branch was the worst of it. Below 0.80, Jev was almost always wrong about what kinds of documents a post discussed.

So we set a per-branch rule we called Policy A. Anything Jev scored at 0.90 or higher was kept automatically. A middle band went to Claude for review, and that band was different for each branch (0.50 to 0.89 for Legal Use Cases, 0.60 to 0.89 for Service, 0.80 to 0.89 for Documents, 0.30 to 0.89 for Area of Law). Everything else was dropped. The sample said that would give us about 94% precision while losing about 18% of the tags a full review would have found.

The full run, by the numbers:

  • Jev made 241,498 yes/no decisions across 14,657 calls on all 2,269 posts. Total cost: about $3.35.
  • Jev settled about 92% of those decisions on its own: 219,970 answers under 0.30 were dropped, and 1,536 answers at 0.90 or higher were kept (before the parent and cap rules trimmed that to 1,174).
  • 6,336 decisions (about 2.6%) went to Claude for review, spread over 1,347 posts. Claude kept 2,118 of them, and 1,910 survived the parent and cap rules.
  • Counting the calibration sample, Claude made 7,946 review calls on the SALI run, handled by 34 review agents running in batches.
  • Final result: 3,431 SALI tags on 1,151 posts. Legal Use Cases carried the load with 2,569 tags, followed by Document / Artifact (443), Service (220), and Area of Law (199).

The most common tags were Business of Law, Legal Research, Pricing Management, Knowledge Management, and Legal Process Improvement, which sounds about right for a blog that has spent eighteen years arguing about billable hours and law libraries.

Run 2: CLOC Core 12

CLOC’s Core 12 is a flat list of twelve legal operations functions (Financial Management, Knowledge Management, Technology, Training & Development, and so on). With no tree to walk, each post got exactly one Jev call with twelve questions.

The calibration step used the same 100 posts, with 318 judgments this time. Jev behaved much better on the flat list:

Jev’s score How often it was right
0.30 to 0.49 16%
0.50 to 0.69 41%
0.70 to 0.89 69%
0.90 and up 100% (47 of 47)

We auto-kept everything at 0.90 or higher and reviewed 0.50 to 0.89. For Information Governance and Business Intelligence we reviewed down to 0.30, because the sample showed real tags hiding lower down. The expected loss was about 7%.

The full run, by the numbers:

  • Jev made 27,228 yes/no decisions in 2,269 calls, in about two minutes, for about $0.39.
  • 450 decisions were 0.90 or higher and kept automatically (403 after the sample posts were swapped for their reviewed versions). 22,682 were under 0.30 and dropped.
  • 2,104 decisions (about 7.7%) went to Claude across 1,166 posts. Claude kept 1,298.
  • Counting calibration, Claude made 2,422 review calls on the CLOC run, handled by 13 review agents.
  • Final result: 1,849 CLOC tags on 1,190 posts. Technology led with 542, then Financial Management (214), Training & Development (149), and Knowledge Management (145).

What Worked

Letting Jev do the boring part. Across both runs, Jev answered 268,726 yes/no questions for under four dollars. Most of those answers were obvious “no” calls (a 2009 post about Twitter search tips is not about Bankruptcy Law), and Jev knocked them out at a quarter-second a call. That’s exactly the kind of job Nate described in his post, and on this project it held up well.

The quote rule. Requiring a verbatim quote for every kept tag turned out to be the single best quality control in the project. A script checked all of those quotes against the source text, and zero failed. It also means that every reviewed tag in the spreadsheet comes with its own receipt, which the librarian in me loves.

Calibrating before scaling. The 100-post sample cost thirty cents in Jev calls and one afternoon of review, and it kept me from shipping a dataset where a third of the tags were wrong.

Running Claude review inside Cowork. Moving the review out of the paid API and into my regular session changed the cost picture completely. The API version would have been something like $70 to $110 for the SALI run alone.

CLOC fit the content. A flat, twelve-item list turned out to be much easier for Jev than a deep tree, and it maps better to what this blog is actually about. Knowledge Management shows up on 145 posts under CLOC, compared with 83 under SALI.

Transcripts held up. I expected the long, rambling podcast conversations to confuse things. They didn’t. Transcript posts and regular posts scored about the same (55% and 53% precision at 0.70 and up in the SALI sample).

What Was Questionable

SALI is the wrong shape for this corpus. SALI was built to describe legal matters. Area of Law barely registered (199 tags across 18 years), and most of the work landed in Legal Use Cases. That’s a reasonable outcome for a business-of-law blog, but it tells me SALI is a better fit for matter data and client documents than for commentary about the industry.

Nearly half the posts have no SALI tags. 1,118 posts came out with no SALI tags, and 805 have no tags from either taxonomy. Plenty of those deserve to be empty (I’m not sure “Merry Christmas!” or “The Leadership and Management Styles of Dolly Parton” needs a SALI tag), but some are real misses. The AALL 2026 Annual Meeting preview should have picked up Legal Research, and it didn’t.

The reviewers haven’t been reviewed. All 10,368 review calls were made by Claude. Nobody has graded Claude against a human yet, and that includes me. The quote rule tells me the evidence really is in the post, and I still need a person to tell me whether Claude read it the right way.

The 84% auto-keep rate on SALI. Tags Jev scored 0.90 or higher went straight in without review. On the sample that was right 84% of the time, so something like one in six of those 1,174 auto-kept SALI tags is probably wrong.

Jev’s confident mistakes. Jev said a confident “yes” to Expert Opinion and Template on the Best Lawyers episode, and to Opinion Memo Practice on the LexisNexis CTO interview. That’s why the per-branch thresholds exist. Nate made this point in his post as well: calibration is measured across thousands of answers, so a 0.95 on any single post is still just a strong hunch.

What happened below 0.30. We never reviewed anything Jev scored under 0.30, so whatever real tags are hiding down there are unmeasured.

What Failed

The first dollar. The first pilot sent Jev’s close calls to Claude through the paid API. Claude kept all 85 of the borderline tags it reviewed, which makes a review step pretty pointless. Then 17 episodes failed after Claude had already answered and been billed. Fifteen of those failures happened because a fix I thought had been saved to my computer never actually landed, so the old script kept running. To top it off, the script didn’t count the cost of failed episodes, so my $1 key limit was gone before the report said it should be. We fixed the logging, tightened the prompt, started checking every saved file with a checksum, and eventually moved the review into Cowork.

The RSS feed was the wrong source. The podcast feed cuts show notes off at about 4,000 characters, right where the “Transcript” heading starts. Everything on the first podcast run was tagged from summaries only. The blog feed has the full text, which is why we switched.

I knocked over my own blog. While downloading the archive, about 65 pages in, geeklawblog.com started throwing Cloudflare 520 and 521 errors. My back-to-back requests probably didn’t help. We waited for it to come back and pulled the rest slowly, one page every few seconds.

Sponsor reads. About half the recent transcripts open with a Legal Technology Hub segment on AI governance, which would have tagged every episode with AI governance. The first attempt to strip it out missed most of them. A second pass that looked for the sponsor’s guest speaker did better, but it’s still a rough heuristic.

Where This Leaves Me

For less than five dollars in Jev and OpenRouter charges, plus about 47 review agents’ worth of Cowork usage, every post on the blog now has two sets of standard tags. Every tag comes with a Jev score, a record of who made the call, and in most cases a quote from the post itself. That’s a catalog I would have killed for back when I was running a law firm library with a card catalog and a budget line for “cataloging backlog.”

The part that sticks with me is how much the calibration step mattered. I went in assuming Jev’s confidence scores could be taken at face value, and on a deep taxonomy like SALI that assumption would have burned me. Thirty cents and 100 posts worth of spot-checking showed me which score ranges needed a second reader.

This was a fun initial test that presented a few areas that need improvement, but there is a lot of potential here to open up a number of decision processes at a very affordable price.