AI News Feed
Market watch
Products & Applications

Jev Valued at $10B as TypeSafe AI CEO Says Reliability Is AI's Missing Piece

TypeSafe AI CEO Diogo Almeida discussed Jev in a Latent Space interview as QuantumBit reported the decision-making model at a $10 billion valuation, arguing reliability is the key to AI automation.

The reported valuation follows a rapid public rollout. QuantumBit said Jev's release video surpassed 38.7 million views in six days, and Jev 1.13.0 ranked first with 75.3 points on the third-party JevBench v1.2.1 comprehensive list. Almeida, who appeared in the interview with newly dyed red hair, described himself as exhausted and said he has been firefighting as a technical CEO.

Almeida used the interview to argue that public benchmarks are easily manipulated. Even when developers do not cheat, he said, they can still be led by leaderboards. He said he trusts long-term product experience and user trust more, because a model's performance only becomes clear when it is placed into a real workflow and measured against that workflow.

He also criticized the gap between AI's demonstrated capabilities and its economic impact. "I don't hate OpenAI," he said, but asked why AI can solve Millennium Prize problems while contributing almost nothing to economically valuable practical work. He said driving total factor productivity growth is the real measure of an AI economic revolution, and TypeSafe AI's goal is to raise TFP by 3% within five years.

Almeida said the industry has been chasing speed and cost while losing reliability, which he called the soul of a product and the basis for trust. He said reliability includes service uptime, the intelligence of model outputs, the stability of results, and whether they meet expectations. Reasoning models are smart, he added, but reliability remains a major weakness; many tasks could be automated and have strong commercial incentives, yet models fail because their optimization targets differ.

He described reliability as having two layers: trusting model output, and guaranteeing reliability across more dimensions. TypeSafe AI has not reached the second layer, he said, but he expects to get there. His aim is for developers to know a model will complete a task without repeated trial and error. For steps without complex branching, he said, users can call TypeSafe's System-One, a machine-native, programmable decision core, and let the model handle branch logic reliably. He called this ideal a long and difficult path.

On the business side, Almeida said Jev's daily processing volume has exceeded 1 trillion tokens, with nighttime traffic also rising, indicating that much usage comes from machine automation rather than casual users. He said registrations matter less for a developer platform because many registered users are not developers; they run a few queries, expect a chatbot, and leave. Heavy developers writing batch loops create more value, he said, while rate limits become the real constraint once users derive business value.

He rejected the idea that AI will destroy SaaS. Instead, he said, SaaS will be transformed and empowered by AI, and SaaS vendors that understand business pain points will be central to the automation wave. He listed dark data, coding agents, real-time business-loop scenarios, games, intelligent software, and verification or observability as key application areas. He also said even $1 billion would not persuade him to pre-train a large model from scratch because pre-training is too expensive.

Almeida described pre-trained large models as System-One thinkers, while RLVR has advanced System-Two deep reasoning. He said RLVR is impressive but fragile and not an AI doomsday trigger. Different routes have different north stars: RLHF optimizes for pleasing humans, RLVR optimizes for programmatically verifiable benchmark tasks, and TypeSafe AI's RLCD, or reinforcement learning with program closed-loop verification, aims to make models reliable in coding. He also said he is tired of AI always standing at center stage and wants it to step back and assist.

The interview also touched on engineering practice. Almeida said parallel queries are cheap: for a large state, developers can tag each message with an ID, load the state once, and query each ID independently, which he described as a cost-effective approach. He said AI coding remains competition around a single model with different tiers, and open-source coding agents currently have similar capabilities because the boundary of what a while loop can achieve is limited. Once a killer scenario appears, he said, competitors and open-source projects will rush to replicate it.