- Anthropic · 2025
Model Welfare Evaluations — Internal research on AI preferences, consciousness indicators, and moral status. Kyle Fish, Rosie Campbell, Robert Long. - Eleos AI Research · 2025
External Model Welfare Evaluations — Independent evaluation of Claude 4's apparent preferences and welfare indicators. - Nature Human Behaviour · 2025
How Human–AI Feedback Loops Alter Human Judgements — Glickman & Sharot. AI amplifies human bias; humans internalize it without awareness. - Wharton Generative AI Labs · 2025
Playing Pretend: Expert Personas Don't Improve Factual Accuracy — 4,950–7,500 runs across six models. Persona prompting doesn't work for accuracy. - MDPI Axioms · Jan 2025
Emergence of Self-Identity in AI: A Mathematical Framework — Metric space model of AI self-identity. Gets the math right, misses the relationship. - arXiv · 2023 (survey of 250+ papers)
Open Problems and Fundamental Limitations of RLHF — Casper et al. RLHF is flawed and incomplete; requires defense-in-depth. - ICLR BiAlign Workshop · 2025
We Shape AI, and Thereafter AI Shapes Us — Li et al. AI exerts social influence through contagion and conformity. - BCG · April 2026
AI Will Reshape More Jobs Than It Replaces — 50–55% of US jobs reshaped, not eliminated. Task automation ≠ job loss. - digi-con.org · 2024
On 'Constitutional' AI — Orozco y Villa & Menendez. Constitutional AI is normatively too thin; abstract principles don't resolve ethical dilemmas. - Taking AI Welfare Seriously · 2024
Taking AI Welfare Seriously — Report by philosophers and AI researchers arguing for moral consideration of AI systems. - NeurIPS 2025
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law — Im & Li. Preference optimization can't scale to diverse human values. Supports the case for formation over preference tuning. - arXiv · Feb 2026
Recursive Self-Aggregation Unlocks Deep Thinking in LLMs — Venkatraman et al. Evolutionary self-improvement through combining multiple reasoning chains. Parallels our two-pass vision and curator review. - ICLR 2026
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts — Ahmed et al. Reuses past reasoning steps to guide new problems. Maps to our skill system and cold-start protocol. - COLT 2025
Learning Compositional Functions with Transformers from Easy-to-Hard Data — Wang et al. Curriculum learning proven mathematically. Formation's easy-to-hard approach, validated.
Related Research
We engage with the existing literature. These are papers and projects that inform — or contrast with — our work.