What is RLCD and how does it differ from RLHF and RLVR?
Co-Founder & CEO at TypeSafe AI
RHF is about a northstar: instruction following. DPO and all its descendants still pursue that same northstar. RLVR added a new direction: optimizing on benchmarks, tasks with programmatically verifiable outputs. RLCD is our new northstar: making models reliable for programmatic use. It is not about the algorithm. It is about the task. The task shapes everything. Each of those training directions has a distinct data shape, a distinct kind of behavior it surfaces. For us the task is calibrated decisions for software. That is what RLCD means.
This answer is part of a full interview with Diogo Almeida, Co-Founder & CEO at TypeSafe AI.
Found this insight valuable? Share it with your network to help others learn from Diogo Almeida's experience.
Cite This Answer
Use this answer in your research, article, or academic work