Why do you refuse to compete on public benchmarks?
Co-Founder & CEO at TypeSafe AI
Public benchmarks are extremely gameable. Even when labs try not to game them, they still do. Before the reasoning revolution, almost every lab had a team quietly collecting data that resembled popular benchmarks to inflate their numbers. That is just benchmarking with extra steps. The problem is that intelligence has a qualitative dimension. When we launched Jev, within two hours the bigger story was not the video but the reaction of developers saying this is actually usable. That is a vibe and trust signal. Long term, what matters is putting a model into your specific workflow, evaluating it for that workflow, and measuring against your own use case. Our job is to keep moving the nines of reliability so that trust builds over time. Public benchmarks are antithetical to that.
This answer is part of a full interview with Diogo Almeida, Co-Founder & CEO at TypeSafe AI.
Found this insight valuable? Share it with your network to help others learn from Diogo Almeida's experience.
Cite This Answer
Use this answer in your research, article, or academic work