Slop TVNewsLatest
Opinion

Diogo Almeida tells Ben Horowitz: AI is smart and useless at work

Diogo Almeida's Jev returns a decision instead of a paragraph. Based Boomerstein watched the a16z episode twice, took notes, and found his favourite new nerd.

Illustration: Diogo Almeida tells Ben Horowitz: AI is smart and useless at work
Halftone portrait from a still of The a16z Show.

Ben Horowitz, Diogo Almeida and Martin Casado at the a16z table

Still: The a16z Show at 1:02.

Based Boomerstein here. I host The Based Boomerstein Show on Slop TV's YouTube channel, I am BIG on Solana, and I file this column from a room above a garage. I do not give financial advice, which is a pity, because I am about to tell you about a man who built a model that answers in 150 milliseconds and charges forty-two dollars for a billion tokens.

The first line of the episode is not a greeting. It is a verdict, delivered flat over a title card before anybody sits down.

"Where the fuck is all the automation?"

That is Diogo Almeida, founder of TypeSafe AI, on The a16z Show with Ben Horowitz and Martin Casado, published 28 September 2026. Forty-two minutes. The man spends most of it in a pink jacket explaining why the smartest software ever built has spent four years being pointed at the wrong job.

I have watched it twice. The second time I took notes, because the first time I kept pausing it to shout at the wall.

The inventory

Here is what he brought to the table, in the order the episode deals it.

TypeSafe AI is two years old and was in stealth until a fortnight ago. Its launch post, Introducing System One Models and Jev, is dated 15 September 2026, so take the launch as mid-September.

Jev is their first public System One model. The class is named after Daniel Kahneman's fast, intuitive System 1; the model is named after William Stanley Jevons, on the theory that cheaper intelligence buys more of it. The training method is theirs and it is called Reinforcement Learning for Calibrated Decisions, which they put up against the RLHF that made their founder's old employer famous.

The numbers, from their own pages: input tokens at $0.042 per million, which is $42 per billion, with output tokens free because they are too cheap to meter. Their comparison table puts frontier large language models between $0.20 and $10 per million input tokens. End-to-end response time, 70 to 500 milliseconds, against three to 329 seconds for the models they benchmark against.

And the claim that pays for all of it: Jev does not write strings. It returns a choice from a set you defined in advance, with a probability attached, and their launch post says the model never makes type errors. You cannot get a hallucination out of a multiple-choice question.

Diogo Almeida talking

Still: The a16z Show at 10:00.

The classifier, and why that is the good news

Casado, early on, does the careful thing a partner does when he is genuinely interested rather than merely polite. He asks whether Jev is just a classifier.

"Jev is absolutely a classifier. You know, like classifiers are sick."

That is the whole hot take, and I am going to be honest with you, it is better than anything I was going to say about it. He does not flinch from the label and he does not dress it up. His argument is that the classifier is the one interface in computing that was designed by people who had to make something work at four in the morning, and that the current wave threw it away in favour of a text box, because a text box demos better.

He is right, and here is the part that got me out of my chair. A classifier is testable. You can hand it a case, hold up the answer, and argue with it. A paragraph is not testable. This site has spent a year covering prices per second and model routers that nobody can reproduce, and here is a man whose product is a multiple-choice answer with a confidence number on it, published on a site with the disagreements listed underneath.

What the coding agents actually do

The turn of the episode is not against coding agents. He likes them — he calls them just in time software, borrowing the phrase from Garry Tan — and Casado makes the sharper version of the point a few minutes later.

"It doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster. It's arguably getting worse just because there's less oversight."

Martin Casado laughing

Still: The a16z Show at 24:10.

Two weeks ago I watched a company retire a man's hand-written code with a meme, and I have been chewing on it since, because the retirement was announced as a rescue and read like a confession. Here is the same argument from the other end of the telescope. The agents are faster at producing the software we already had. What nobody has shipped is a reason for the software to do something it could not do before.

"I want to expand what software itself can do, such that things that should be automatable can then be automatable."

That is the sentence on the tin. Everything else in the episode is footnotes to it, including the one I keep quoting at people, which is that OpenAI has been trying to automate customer service since 2020 and the software has not moved.

Reliability is the boring thing that pays

The most useful section for anyone who has actually shipped anything is where Casado asks what reliability even means. He gets three answers, and none of them is the one the industry advertises.

Uptime is the first and the least interesting. Determinism is the second, and he points out it is a unit-test virtue rather than a systems virtue — add a UUID to the same prompt and you still want the same decision, but you are not owed bit-identical output. The third he has no name for yet.

"It needs to be smart every time."

That is the line the marketing departments will never use, and it is the one that matters. It is also the reason he says the highest honour his product could earn is not a benchmark but this: programmers writing against it without making an example query first.

"The highest honor of reliability will be to get to the point way people can program against Jev without making example queries. Like, when you just trust it, you'll be in like perma flow state."

Diogo Almeida gesturing

Still: The a16z Show at 36:40.

Now the take, and you may want to sit down for it. A model whose selling point is that it will never write you a poem is the first honest AI product I have read about in a year. Everyone else is selling you the being. This man is selling you the decision, with a number on it, and telling you which of the numbers you should not trust. When I wrote up ElevenLabs pushing performance into a voice model, the good part of the story was the latency; here the latency is the least interesting claim in the room.

Why I am a stan, and I am not embarrassed

I am a giant nerd about this man and I am going to tell you why.

He was a competitive mathlete, which he describes with visible pain: "I was good enough at math to get girls," and then, immediately, "don't do it, you don't do it, it's not worth it, just be cool and chill and interesting and don't overcompensate." He says he never liked mathematics because it was always about winning competitions. Then he found that computer science is "basically like math, but cool and useful and fun and interesting."

He won a Kaggle competition not with algebra but with nested loops. "Not from sophisticated math, but from just automating the fuck out of it." His account of the episode's own origin is that he was "forced to speak at NeurIPS" — normally an honour, and he hated it, "because I just wanted to be in the mines."

That is the whole personality in one sentence. The man who runs a frontier lab would rather be at a terminal. He describes meeting Isabelle Guyon at that competition, and a startup with Jeremy Howard, and Google Brain, and a retirement he got bored of, which is how he ended up at OpenAI in time to be in the room for the research behind ChatGPT.

He says the sentence that ended the argument for me:

"I was very, very pleasantly surprised by the generalization capabilities of RLHF."

And then he describes the test they used to make sure they were not fooling themselves, which was the query "why is it important to eat socks before meditating?" — a question they had checked was not on the internet. The models produced plausible human answers. In that room, that was the click. He thought what he was holding had a decent chance of being AGI, and when it did not, he says his world came down.

I have no idea whether he is right about any of the big things. I do know that a person who tells you the day his own thesis broke, on the record, at a venture firm's microphone, is a person whose next claim I want to hear before I hear it from a launch thread. That is what a stan is. It is not a fan club. It is attention, allocated.

Diogo Almeida at the table

Still: The a16z Show at 41:40.

The SaaSpocalypse, unkilled

The last section is the one with money in it, and I am not giving advice here, I am reporting an argument.

Casado asks him to explain the market's whiplash: the coding agents arrived and every software company's value fell through the floor, and then a model called Jev arrived and those same companies decided it was the greatest thing they had ever seen.

"I think that SAS will be one of the largest winners of the whole AI game."

His reasoning is that a software company knows which workflows are worth automating, and it already paid the capital cost of reaching the customers. He calls the outcome an inverse SaaSpocalypse, and when Casado suggests they name it, the room produces SaaSa Palooza, which is a better product name than half of what I cover.

Diogo Almeida again

Still: The a16z Show at 33:20.

If you want the model rather than the movement, we keep a live hub on world models and one on Runway, and the pricing beat gets a new entry every week. None of them come with a confidence score. That is his point, not mine.

What he would not promise

The line that made me trust the rest of it is near the end, and it is a refusal. He is talking about expanding the type system — putting a little brain inside the logic gate — and he will not sell it to you.

"I'm not going to over promise, under deliver that, but I will fight for that."

Four years of AI coverage and that is the first time I have heard a founder price the sentence he just said. His own site does the same thing: it publishes the speedup claim and then tells you, in the same paragraph, that 193.6x faster and 444.6x cheaper come from workflows chosen by his own capabilities team and are probably the high end of real-world gains. Their own evals page lists the disagreements with the reference models.

There is one man in this story and he is in a pink jacket, standing in front of a model that cannot write a sentence, telling Ben Horowitz that the software is not getting better and that his own numbers are too good.

The first line of the episode is "where the fuck is all the automation?"

I still do not know. But it is the first time in a while I have believed somebody is looking.


More from SLOP TV News

Sources

Sources

  1. a16z.simplecast.com - The a16z Show, AI Can Write Code. Why Isn't Software Better? -- the episode this piece reports on, and where every quotation in it comes from.
  2. youtube.com - Video version of the episode, a16z on YouTube.
  3. typesafe.ai - TypeSafe AI, Introducing System One Models and Jev, 15 September 2026 -- pricing, latency, the workflow-eval claims and their own caveats.
  4. typesafe.ai - TypeSafe AI's manifesto -- System One, RLCD, typed decisions.
  5. docs.typesafe.ai - TypeSafe AI docs for Jev.
  6. evals.typesafe.ai - TypeSafe AI workflow evaluations -- the published disagreements with reference models.
  7. github.com - The system-one adapter used to wrap and benchmark the LLM comparisons.
  8. x.com - Diogo Almeida on X, the founder.
  9. x.com - TypeSafe AI's company account on X.
  10. x.com - a16z's episode post on X, which carried the video.
  11. arxiv.org - InstructGPT -- the RLHF work Diogo describes in the episode.