Shi Feng
I’m an assistant professor leading a research group on AI alignment. Our long-term goal is to ensure human oversight and control over future AIs (motivation explained in this talk). To get in touch, consider using the contact form instead of email.
I’m recruiting:
- MATS scholars for summer 2026
- PhD students starting fall 2026 and spring 2027
- Research scientists / postdocs to start anytime
- Collaborators to start anytime via Praxis Research
My current focuses are:
Deception
- Language Models Learn to Mislead Humans via RLHF
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
Collusion
- LLM Evaluators Recognize and Favor Their Own Generations
- Spontaneous Reward Hacking in Iterative Self-Refinement