How do you think about the agent alignment problem at the personal level, not the civilizational level?
Founder & CEO at Instinct
If you hire an employee, you pay them and that creates aligned incentive. There is no other revenue stream driving their behavior. The question with an AI agent is what is its primary objective. Instinct follows higher level objectives rather than just being a task accomplisher. A task accomplisher takes whatever the user says literally and executes it. We instead trained Instinct around objectives like build genuine trust with the user, watch over the user, have their back when things are dropped, make the user feel genuinely safer. When the user asks for something well-intentioned, the way to demonstrate trustworthiness is simply to do it well. That higher level objective framing also makes the system more robust to edge cases and adversarial inputs because the model is optimizing for your genuine wellbeing rather than just literal instruction following.
This answer is part of a full interview with Noah Shinn, Founder & CEO at Instinct.
Found this insight valuable? Share it with your network to help others learn from Noah Shinn's experience.
Cite This Answer
Use this answer in your research, article, or academic work