Register
F
Professor profile

Fazl Barez

University of Oxford · Mechanical Engineering
Technical AI safety alignment Mechanistic interpretability neural network analysis Internal representations reasoning in large language models Detection mitigation of harmful or deceptive model behaviour

About
Regular biography

Fazl Barez is a Senior Researcher at the University of Oxford, affiliated with the Department of ME. His research focuses on technical AI safety and alignment, mechanistic interpretability, and neural network analysis. He investigates internal representations and reasoning in large language models, as well as methods for detecting and mitigating harmful or deceptive model behavior. Barez also works on verification and evaluation of model reasoning processes, robust removal of dangerous capabilities from AI systems, and automated interpretability. He is affiliated with the Cambridge's Centre for the Study of Existential Risk, ELLIS, and other international research centres. At Oxford, he teaches the AI Safety and Alignment course and supervises students in AI safety and machine learning systems.


Scholar profile summary
Scholar-generated biography

Fazl Barez is a researcher at the University of Oxford with expertise in Machine Learning, AI Safety, Explainability, Interpretability, and AI Governance and Policy. His work explores critical challenges in AI systems, including deceptive behaviors in large language models, reward tampering, and the limitations of chain-of-thought reasoning as a form of explainability. Barez also investigates methods to improve the interpretability of vision-language models and assess vulnerabilities such as data poisoning and hallucinations in LLMs. His research emphasizes the development of rigorous benchmarks and best practices for AI safety and reliability.

Source: google_scholar · 90 words
Related professors