CoRe Safe AI @ CMU

Publications

Representative work by CoRe Safe AI members

Against Proxy Optimization

Sven Neth

Philosophy and Phenomenological Research · 2026 · Journal · arXiv

Core Safety Values for Provably Corrigible Agents

Aran Nayebi

AAAI 2026 Workshop on Machine Ethics · 2026 · arXiv

A Correspondence Problem for Mathematical Proof

Simon DeDeo and Eamon Duede

Philosophy of Science (forthcoming) · 2026 · arXiv · PhilArchive

Designing Rules to Pick a Rule: Aggregation by Consistency

Ratip Emin Berker, Ben Armstrong, Vincent Conitzer, and Nihar B. Shah

ICLR 2026 · 2026 · OpenReview · arXiv

AI Testing Should Account for Sophisticated Strategic Behaviour

Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, and Vincent Conitzer

NeurIPS 2025 · 2025 · arXiv · OpenReview

Longtermist Myopia

Amanda Askell and Sven Neth

In Essays on Longtermism, Oxford University Press · 2025 · OUP · PhilPapers

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

Ziming Luo, Atoosa Kasirzadeh, and Nihar B. Shah

NeurIPS 2025 AI for Science Workshop (Spotlight) · 2025 · arXiv

Off-Switching Not Guaranteed

Sven Neth

Philosophical Studies · 2025 · Journal · PhilPapers

AlephZero and Mathematical Experience

Simon DeDeo

Bulletin of the American Mathematical Society · 2024 · Journal · arXiv

Assessing Group Fairness with Social Welfare Optimization

Violet Chen, J. N. Hooker, and Derek Leben

CPAIOR 2024 · 2024 · Springer · arXiv

Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, and William S. Zwicker

ICML 2024 (Position Paper) · 2024 · PMLR · arXiv

A Dilemma for Solomonoff Prediction

Sven Neth

Philosophy of Science · 2023 · Journal · PhilSci-Archive

Foundations of Cooperative AI

Vincent Conitzer and Caspar Oesterheld

AAAI-23 (Senior Member Presentation Track) · 2023 · AAAI · PDF

The Tragedy of the AI Commons

Travis LaCroix and Aydin Mohseni

Synthese · 2022 · Journal · arXiv

Ethics for Robots: How to Design a Moral Algorithm

Derek Leben

Routledge (book) · 2018 · Publisher · PhilPapers