Diogo Almeida, Co-Founder & CEO at TypeSafe AI
4/5 Rating
AI
Pre-revenue/mo
Pre-Revenue ARR

Diogo AlmeidaCo-Founder & CEO

In this interview, Diogo Almeida, co-founder and CEO of TypeSafe AI, breaks down why the industry has been building the wrong kind of AI for years and how Jev, TypeSafe's new system one model, offers a fundamentally different approach. He explains RLCD, the dangers of mode collapse in RLHF, why calibration beats benchmarks as a measure of intelligence, his vision for AI disappearing into the background of software, and why the real AI economic revolution has barely begun.

Diogo Almeida

Diogo Almeida

Co-Founder & CEO

TypeSafe AI

TypeSafe AI

Founder Stats

  • AI
  • Started 2024
  • Pre-revenue/mo
  • 50+ team
  • San Francisco, California, USA

About Diogo Almeida

Diogo Almeida is the co-founder and CEO of TypeSafe AI, a San Francisco-based AI lab that emerged from stealth in September 2026 with 40 million USD in seed funding. A former researcher at OpenAI and Google Brain, Almeida is a co-creator of ChatGPT and a co-author of the landmark InstructGPT paper that established reinforcement learning from human feedback as an industry standard. He founded TypeSafe with CTO Erik Gafni and COO Sasha Sheng to pursue a fundamentally different north star for AI: machine-native, programmable intelligence designed to run deep inside software systems rather than chat with humans.

Interview

September 28, 2026

Q

What is Jev and how is it different from existing AI models?

Question 1 of 14
Diogo Almeida

We needed a new class of models. The most accurate name we have come up with is system one models. The goal is for code to be the consumer. As opposed to large language models which are meant for autocomplete of the internet, or RLHF models which are meant to reply to text, Jev is designed to have its outputs directly consumed by code. Everything we do beyond the internals of the model is optimized for software. Jev is our first large programmable model. It is optimized for intelligence per dollar, which is where the name comes from, a reference to Jevons paradox. We are going all out on that tradeoff.

0
Q

Why is mode collapse in RLHF such a serious problem that almost nobody talks about?

Question 2 of 14
Diogo Almeida

The spicy take is that I believe Yann LeCun is among the closest to reality in his criticisms of large language models. He has this famous slide about how LLMs are doomed as sequence length increases because the probability of an error grows monotonically. Mathematically obvious, yet empirically it does not hold. The reason is mode collapse. When you train with RLHF, the model stops covering the full distribution of possible answers and drops the minority cases to concentrate on what gets rewarded. This is exactly what GANs did when they started making blurry images instead of diverse ones. The model becomes hyperconfident. It needs to be miscalibrated in order to not go off the rails because errors are very easy to detect and subtle wrongs are very hard to spot. That miscalibration then warps the entire probability space of the model.

0
Q

What is RLCD and how does it differ from RLHF and RLVR?

Question 3 of 14
Diogo Almeida

RHF is about a northstar: instruction following. DPO and all its descendants still pursue that same northstar. RLVR added a new direction: optimizing on benchmarks, tasks with programmatically verifiable outputs. RLCD is our new northstar: making models reliable for programmatic use. It is not about the algorithm. It is about the task. The task shapes everything. Each of those training directions has a distinct data shape, a distinct kind of behavior it surfaces. For us the task is calibrated decisions for software. That is what RLCD means.

0
Q

Why do you refuse to compete on public benchmarks?

Question 4 of 14
Diogo Almeida

Public benchmarks are extremely gameable. Even when labs try not to game them, they still do. Before the reasoning revolution, almost every lab had a team quietly collecting data that resembled popular benchmarks to inflate their numbers. That is just benchmarking with extra steps. The problem is that intelligence has a qualitative dimension. When we launched Jev, within two hours the bigger story was not the video but the reaction of developers saying this is actually usable. That is a vibe and trust signal. Long term, what matters is putting a model into your specific workflow, evaluating it for that workflow, and measuring against your own use case. Our job is to keep moving the nines of reliability so that trust builds over time. Public benchmarks are antithetical to that.

0
Q

You are against safety alignment in APIs. What is your actual position?

Question 5 of 14
Diogo Almeida

I am not opposed to safety as a principle. My position is that safety alignment is misaligned with users when applied to an API. For a first party product like ChatGPT, it makes complete sense. Maybe that company does not want to support certain types of content for their user base. That is fine. But in an API, a refusal is a type error. If you are a developer and your code dependency randomly refuses a query because a user sent a weird message, your software stochastically breaks. That is not a safety feature. That is an antipattern that comes from people who do not understand software. Refusals fracture the intelligence of the model every single time they are baked in. And every single time you overfit the model to some weird edge case to prevent refusals, you fracture it more.

0
Q

How do you think about AI safety at the foundation model layer?

Question 6 of 14
Diogo Almeida

Would I prefer our technology not be used to harm people? Obviously. Would I put my thumb in the scale for that? Yes, in some ways. But will I do it at the technological layer? Absolutely not. I think intelligence will be more like a database than a coworker. We do not ask databases to police their own downstream use. The first line of defense is that you are liable for what you build with any tool. We should give developers maximum power and trust that they are building things for good purposes. We obviously cannot even know what downstream tasks our API is being used for since it should be decomposed into small, isolated calls. That is a healthy boundary to maintain.

0
Q

Why do you place so much emphasis on calibration over raw accuracy?

Question 7 of 14
Diogo Almeida

Calibration means that when the model says 80 percent confidence, it is actually right 80 percent of the time. This matters enormously for software because thresholding is how you control behavior. If the model is miscalibrated, your threshold is meaningless. You want similar inputs to produce similar outputs. We test this by putting unique IDs called nonces into prompts that are semantically identical and checking that the outputs are consistent. This is robustness, and it is where AI most frequently burns developers today. They think the model should be smart enough to automate a task, there is economic incentive to automate it, yet they cannot trust the output because the distribution is too jagged.

0
Q

What is your argument against function calling as the standard API interface for AI decisions?

Question 8 of 14
Diogo Almeida

Function calling is a deeply antipattern interface for AI decisions. One very simple ask I always had was that for any function call you should get a calibrated probability for each option. Without that, the only way to control refusal thresholds, routing thresholds, or decision confidence is through begging in a system prompt. That is an insane interface for software engineers. Our three primitives, noul for boolean decisions, choice for structured selection like a switch statement on an enum, and score for sorting or thresholding, map directly to programming primitives. They are not new names for existing types. They are new concepts that give software engineers the ability to actually control and measure AI behavior programmatically.

0
Q

How should developers decompose tasks when building with system one models?

Question 9 of 14
Diogo Almeida

Break everything down into its smallest semantic unit. Never ask one big question when you can ask ten small ones. For each piece of state you want to reason about, ask a separate, explicit, very structured question. This way every single call is evaluable, every failure is diagnosable, and every fix is durable because it lives in the structure of your code, not hidden in some context-rotting system prompt. If you find a situation where the model got something wrong, you add a new question to handle that case. The bug is now solved forever. It is like machine learning without the machine learning. And yes, it is more calls, but with a model like Jev the cost is not a barrier and the reliability gain is enormous.

0
Q

What is your vision for where TypeSafe and machine-native AI are heading?

Question 10 of 14
Diogo Almeida

We are going to keep shipping new shapes of intelligence. Jev is the first system one model. That is not the only category. Some things are not decisions. There will be other types that are equally machine native. Longer term, I think about this like the early internet. Think about what composability gave us in the two thousands. Creation was back on the menu. That is what we want to give software engineers again. AI should disappear into the background of software and make everything around you more delightful. The question that haunts me is how can AI solve millennium prize problems in mathematics but still not automate basic repetitive work that does not even take very smart people to do? The answer is the plug does not fit. We are building the plug.

0
Q

Why did you leave OpenAI to start TypeSafe even though you thought a competitor must already be working on this?

Question 11 of 14
Diogo Almeida

When I first thought of this direction, I assumed that it was so obviously correct that Anthropic or someone else must already be pursuing it. I was wrong. I talked to other companies and figured out that starting a startup would be faster than convincing an existing lab to change direction. So we did it. I thought the whole project would take a week. I was unbelievably wrong and I apologize to everyone at OpenAI who I thought I was about to show up. But the reason I could not walk away is that if an AI winter happened and I had not done every possible thing I could to prevent it, I would have considered myself personally responsible, both for what RLHF did to the gap between AI promise and reality, and for not going all in on the direction I believed was where value was actually going to be created.

0
Q

What does it mean for AI development to be driven by the right task?

Question 12 of 14
Diogo Almeida

The bitterest lesson in AI is that the task you choose determines everything. It determines what data is valuable, what capabilities get surfaced, what kind of intelligence you can build. Algorithms and compute matter, but data matters more than both, and having the right northstar matters more than data. This has happened maybe two and a half times in large language model history. RLHF established instruction following as a northstar. RLVR made a slight edit toward benchmark optimization. And RLCD is our new task: making AI reliable for programmatic, machine-native use. Most researchers do not spend enough time asking whether the task they are optimizing for is actually the right one. That is the question that everything else depends on.

0
Q

What would you tell a researcher at a frontier lab who believes they have a better direction but cannot get resources to pursue it?

Question 13 of 14
Diogo Almeida

It really depends on why you want to do it. If you want to play around with research, the big labs are genuinely the best place for that. What I would push back on is starting a new lab just for its own sake. Most new labs I have talked to do not have a clear direction. They want money to run experiments. That destroys value. If you truly have a different northstar, a genuine new task you believe in, then absolutely start a company. But do not try to be a lab. Try to solve something real. Be driven by actual problems you want to see fixed in the world. That was the only thing that kept me going through two years of stealth while nobody believed what we were building was real.

0
Q

You said you considered your launch a success even though most people who tried it during the preview did not get it. How do you think about product market fit for infrastructure products?

Question 14 of 14
Diogo Almeida

More than half the people we had play with Jev before launch simply did not understand it. That was very scary for the non-technical members of the team. But we knew from the computational properties alone that this was going to be huge once the right people touched it. We intentionally targeted developers first because developers find the use cases, then everyone else FOSMOs in. The thing we did not fully anticipate was the rate limit problem. Signups are meaningless for a developer API. What matters is when people start hitting rate limits because they want to run it inside production loops in the background. When machines are calling your API continuously through the night, not just people trying queries, that is the real signal. We surpassed a trillion tokens per day within days of launch and it was not slowing down at night.

0

Video Interviews with Diogo Almeida

Jev: Exclusive Deep Dive — with Diogo Almeida, TypeSafe Co-founder & CEO

Jev: Exclusive Deep Dive — with Diogo Almeida, TypeSafe Co-founder & CEO

Jev: Exclusive Deep Dive — with Diogo Almeida, TypeSafe Co-founder & CEO

Jev: System One Models for Prod, not God — Diogo Almeida, CEO TypeSafe AI (Latent Space)

Jev: System One Models for Prod, not God — Diogo Almeida, CEO TypeSafe AI (Latent Space)

TypeSafe AI Jev Deep Dive: Calibrated Decisions for Developers

TypeSafe AI Jev Deep Dive: Calibrated Decisions for Developers

Cite This Interview

Use this interview in your research, article, or academic work

Related Interviews

Anton Oika, Co-Founder & CEO at Lovable
AI

Anton Oika

Co-Founder & CEO at Lovable

Approx. USD 45 Million
150+

In this interview, Anton Osika, co-founder and CEO of Lovable, explains why the era of vibe coding is over and the era of serious AI-powered…

Read Interview
Alex Karp, Co-Founder & CEO at Palantir Technologies
AI

Alex Karp

Co-Founder & CEO at Palantir Technologies

Approx. USD 680 Million
6000+

In this interview, Alex Karp, co-founder and CEO of Palantir Technologies, argues that major AI labs are quietly steering toward government nationalization to escape unlimited…

Read Interview
Ali Ghodsi, Co-Founder & CEO at Databricks
AI

Ali Ghodsi

Co-Founder & CEO at Databricks

Approx. USD 583 Million
12000+

In this interview, Ali Ghodsi, co-founder and CEO of Databricks, shares the unlikely story of how a reluctant professor became one of Silicon Valley's most…

Read Interview
Rahul Vohra, Founder & CEO at Superhuman
AI

Rahul Vohra

Founder & CEO at Superhuman

Approx. USD 30 Million
1200+

In this interview, Rahul Vohra, founder and CEO of Superhuman, unpacks how to build, scale, and sell a startup with clarity and precision. He shares…

Read Interview
Jack Clark, Co-Founder & Head of Policy at Anthropic
AI

Jack Clark

Co-Founder & Head of Policy at Anthropic

Approx. USD 100 Million
600+

In this interview, Anthropic co-founder and head of policy Jack Clark addresses escalating existential and economic concerns surrounding advanced artificial intelligence. Clark explains the critical…

Read Interview
Sauvik Banerjjee, Group CTO & Chairman of Rezolve India at Rezolve AI
AI

Sauvik Banerjjee

Group CTO & Chairman of Rezolve India at Rezolve AI

Approx. USD 35 Million
450+

In this interview, Rezolve AI Group CTO and Chairman of Rezolve India Sauvik Banerjjee discusses career acceleration, organizational leadership, and enterprise artificial intelligence architectures. Banerjjee…

Read Interview