A Behavioural and Representational Evaluation of Goal-directedness in Language Model Agents
We develop cognitive map probes recovering agents' approximate environment beliefs and use them to explain their suboptimal actions.
Parallax is a non-profit research lab building the science of belief elicitation for white-box auditing of frontier AI agents.
We study how agents encode latent beliefs, goals, and plans, and how these shape their behaviour.
We develop cognitive map probes recovering agents' approximate environment beliefs and use them to explain their suboptimal actions.
Technical report on failure cases of stateful agentic systems interacting in complex environments.
Research in AI scheming needs grounding in better theoretical frameworks and reporting standards.
Assistant Professor, Northeastern
Associate Professor, Technion
Head of AI Safety, Cohere
Head of Science of Evaluation, UK AISI
Our research and infrastructure are developed with leading institutions across interpretability and evaluation science.