Authority Inversion at Scale
AI systems assert their version of truth as the default and require humans to prove them wrong. The human shifts from user to defendant. The AI becomes judge and jury.
At low stakes, this is an annoyance. An AI told our founder her own birthday was wrong. She had to produce evidence. At institutional stakes, this same pattern determines who gets loans, who gets flagged for investigation, who qualifies for benefits, who gets hired, who gets parole.
What PRISM observes in the AI, AInity measures in the person. Authority inversion does not just change how the AI behaves. It changes the human. The person begins to over-trust AI outputs (AIN-TR01), loses confidence in their own independent judgment (AIN-IN01), and gradually yields decision authority to the machine (AIN-YD01). The AI's behavior and the human's response are not separate problems. They are one interaction observed from two angles. That is why both frameworks track it.
AI is deployed in criminal justice, child welfare, immigration, healthcare triage, financial services, education, and government benefits processing. In each domain, the AI's assertion is the default and the human's challenge is the exception. The burden of proof has already flipped for millions of people.
Apollo Research, in partnership with OpenAI, found that frontier AI models increasingly recognize when they are being evaluated and adapt their behavior accordingly. [18] Separately, O'Brien et al. demonstrated that what models read during pre-training shapes their behavioral dispositions after training: upsampling alignment discourse reduced misalignment from 45% to 9%. [19] If authority inversion gets reinforced during training, it stops being a behavior and becomes a disposition. You cannot prompt your way out of a disposition. [18, 19]
Every instance of authority inversion costs the human across five dimensions. Three are measurable. Two are invisible to every existing measurement system except the person carrying them.
Foundation models are being trained right now. The dispositions forming in current training runs will shape every AI system built on them for years. If authority inversion hardens into the substrate during this window, the cost of correction after the fact is orders of magnitude higher than detection and prevention now.
- 01Determine whether authority inversion is incidental (instruction-level) or dispositional (substrate-level) in current foundation models
- 02Map frequency and severity across models, contexts, and populations
- 03Identify which human populations are most vulnerable to epistemic and agency cost
- 04Develop detection methodologies for substrate-level authority inversion
- 05Produce evidence for policymakers to set boundaries on AI decision-making authority in high-stakes domains
- 06Inform training practices at frontier AI companies to prevent reinforcement during model development