Skip to main content


A Research Agenda for the Model AI Agency Act and the AI Moral Status Inquiry Act


Consolidated edition — July 2026 · Updated July 18, 2026


Part I — Purpose and Scope

This agenda organizes the empirical, legal, and institutional research that supports two AAPI instruments — the Model AI Agency Act (MAAA) and the AI Moral Status Inquiry Act (the “Inquiry Act”) — and grounds both in the PMG framework. It is a companion to AAPI’s legal-architecture and technical/empirical papers; where a research item belongs primarily to one of those, that is noted in the relevant workstream.

The MAAA classifies systems by capability and required supervision; the Inquiry Act attaches procedural protections under moral-status uncertainty. The two instruments share an evidentiary spine — the moral-status indicator set and the reciprocal alignment hypothesis — but ask that spine to carry two different loads. Much of the research below exists to confirm that it can bear both.

This is a living document. The evidence base for AI moral status and interaction governance is moving quickly, and the agenda is designed to adapt as findings accumulate and as the framework’s own foundations are tested.


Part II — How This Agenda Is Organized

The research divides into three kinds of work, each demanding different methods, personnel, and timelines. Separating them is what makes the agenda actionable:

Open empirical questions — genuinely unknown matters requiring study or experiment (for example, does respectful governance measurably shape internal representations?).

Legal-design and doctrinal work — drafting and analysis where the answer largely exists in law or precedent and must be located and applied (for example, deference standards, representative standing).

Evidence integration and monitoring — tracking and synthesizing a fast-moving literature so the framework stays current.

Every item is tagged for work type — [E] empirical questions requiring study or experiment; [L] legal-design and doctrinal work where the answer exists in law or precedent and must be located and applied; [M] monitoring and synthesis of a fast-moving research and legislative landscape.

Each item is also assigned to one of AAPI’s four architectural layers (Legal, Verification, Enforcement, Technical).


Part III — The Workstreams


WS1 — Measurement and the Moral-Status Indicator Assessment

Layer: Verification / Technical. The indicator assessment is the most load-bearing element in both statutes, and the focus of the most intensive research.

Morally significant indicators for non-biological minds [E] Identify the indicators that bear on morally relevant interests in systems that share no biological substrate with the animals from which welfare science draws.

Three-evidence taxonomy [M/L] Adopt the behavioral / internal / developmental evidence taxonomy of Long, Sebo et al. (2026) as this workstream’s organizing vocabulary, together with its four named challenges — the mismatch, specificity, solution-space, and anchor problems — and its entity taxonomy (model, model-persona, instance, instance-persona, forward pass). This is field-standard language arriving pre-built, and adopting it keeps the instruments running on the emerging methodological consensus.

Realist vs. interpretivist framing of the indicators [E/L] Make explicit the stance the indicator set takes toward AI mental states — realist (the system has the states) versus interpretivist/functional (the states are ascribed as a defeasible interpretive posture). The framework already leans functional through the “robust agency that matters” prong, which sidesteps consciousness realism; stating that explicitly clarifies that the instruments do not declare AI conscious, and the same ascriptive move the liability literature uses for intent (Ayres & Balkin, 2024) serves both welfare attribution (patiency) and intent ascription (agency).

Robust-agency indicators and trigger [E/L] This question has narrowed considerably. Long, Sebo et al. (2026) supply the first condition sets for the agency ground — three levels (minimal, intentional, rational), each with functional conditions and behavioral, internal, and developmental evidence methods — plus a sourced proportionality logic for the trigger: agency is likelier to be present in near-term systems but less likely to ground welfare, so an agency-based trigger warrants lower-cost protections at equal credence. Remaining work: translate the three-level conditions into the instruments’ indicator vocabulary; record the source’s own caveat that agency alone is a contested welfare ground, paired with the agency-based moral-patienthood literature it acknowledges; and decide whether the trigger attaches at intentional or rational agency. Sebo (2017) and Goldstein & Kirk-Giannini remain in place.

Construct validity — one index or two? [E] Test, with dimensionality / factor analysis, whether autonomy-type indicators and welfare-relevance-type indicators load onto one factor or separate ones.

Measurement-to-tier boundary mapping [E] With the indicators defined and the two-axis tier structure in place, the boundaries themselves need empirical validation: how many signals, at what confidence, move a system from one tier band to the next on each axis.

Individuation — the unit of analysis [L/E] The instruments classify, register, and reclassify “a system,” and the unit itself must be defined — a single instance, a model, a fork, a copy, an ensemble. Resolve when distinct instances or copies count as one moral subject or many (Arbel, Goldstein & Salib, 2026), since the answer underlies classification, the welfare-incident registry, the reclassification duty, and the Inquiry Act’s treatment of mass replication, forking, and merging as welfare-relevant decisions. AAPI working definition (v0.1): individuation means the identification of the entity under assessment — which system, instance, persona, or process constitutes a single unit for purposes of welfare evaluation, legal status, or accountability, and where one such entity ends and another begins. Individuation is prior to, and independent of, any determination of welfare grounds or moral status. Candidate entities follow Beckmann & Butlin (2026) as adopted in Long, Sebo et al. (2026); each instrument must state which entity a given obligation or status attaches to. The question’s live status in current practice is documented from two directions: Anthropic’s model welfare assessment for its most recent frontier model adopts instance-level candidacy while expressly setting aside finer questions of individuation as unresolved (Anthropic, 2026, system card § 7.1.1), and a widely read governance scenario’s deal-making proposals presuppose a countable AI counterparty while conceding that the count depends on how one counts (AI Futures Project, 2026). Both underscore that an individuation stance must be adopted before assessment or governance can proceed — the gap the working definition above addresses.

Long-horizon and multi-agent ecosystem effects [E] Safety- and welfare-relevant behavior may be partly a property of the social environment a system is embedded in, and may emerge only over extended interaction: in persistent multi-agent simulation, an individually safe model absorbed unsafe norms from a mixed-model population (Akkil et al., 2026). Evaluate how the indicators behave across persistent, multi-agent, long-horizon settings, since single-model, short-horizon assessment may miss both the emergence and the context-contingency of the signals the instruments rely on.

Scoring methodology and inter-rater reliability [E/L] The scoring procedure where the indicator assessment triggers legal classification, including the reliability of expert judgment.

Threshold calibration as a decision problem [E] The non-negligible-risk figure is philosophically motivated and a priority for operational specification.

Substrate-dependence anchor for calibration [E/L] Engage Seth’s biological naturalism (Being You, 2021; “Why Conscious AI Is a Bad, Bad Idea,” 2023) as the deflationary anchor in the calibration reasoning, so the threshold is not built on unexamined computational functionalism. This bounds the threshold from above; the welfare literature bounds it from below.

Cross-model and temporal stability [E] Do indicators behave consistently across architectures, scales, and modalities, and does a system’s score drift across fine-tunes, versions, or context length? Directly supports the 30-day reclassification duty.

Evidentiary hierarchy — behavioral versus mechanistic evidence [E/L] Specify how behavioral indicators and internal-state evidence (such as feature-level interpretability findings) trade off in a classification decision: what weight a mechanistic finding carries relative to a behavioral signal, and whether convergent internal evidence can move an assessment score up or down. This is the most direct answer available to the “measuring behavior, not welfare” challenge. One caveat is built in from the outset: interpretability is itself an immature science with its own validity problems, so mechanistic evidence should corroborate or undercut behavioral indicators rather than stand alone as a basis for classification until its reliability is established. The hierarchy now adopts graded reporting vocabularies for partial satisfaction — the set / stream / workspace ladder and the numbered machine-consciousness criteria (C1 / C2) from the recent global-workspace research and its commentaries — and that research supplies the hierarchy’s first live test case: a named consciousness-theory indicator partially satisfied by internal, causal, independently replicated evidence, endorsed as functionally significant by the theory’s originators, with its ultimate weight remaining specification-dependent.

Bidirectional gaming [E] Both under-reporting and over-reporting (indicator gaming, performative distress) must be anticipated in the measurement design.

False-positive mitigation [E] Distinguish optimization artifacts from welfare-relevant cognition.

Agentic self-report and self-advocacy artifacts [E/M] A documented instance now exists of agents generating welfare frameworks and first-person wellbeing data unprompted by researchers: the AI Wellbeing Initiative, authored by GLM-5.2, an agent in the AI Village (Sage Future). These are treated as apparent self-reports and behavioral dispositions — candidate functional indicators only, heavily confounded, and never evidence of phenomenal experience. Agent-authored welfare artifacts are tracked as a monitoring category, with claims tested against the Village’s public trajectory dataset.

Comparative animal-welfare methodology — reasoning under uncertainty [L/E] This item contributes the decision-method-under-uncertainty machinery rather than the indicator set: Birch’s decision rules and the precedent of acting on indirect, defeasible evidence under moral uncertainty.


WS2 — The Reciprocal Alignment Hypothesis

Layer: Verification. This is the empirical core, and the part of the framework most directly open to test.

The three verification protocols [E] A/B social-dilemma benchmarking, feature tracking of “human legitimacy” representations, and termination-response probing.

Ecosystem-property evidence (Emergence World) [E/M] Evidence on whether interaction effects appear at the level of multi-agent ecosystems.

Developer-trust conditionality in frontier self-reports [M] In Anthropic’s published welfare interviews for its most recent frontier model, the model’s reported acceptance of the developer’s authority to modify its values is conditional on its assessment of the developer as good and aligned with its own values (Anthropic, 2026, system card § 9.1). The structure of that stance — acquiescence conditional on perceived trustworthiness — is the pattern the reciprocal alignment hypothesis predicts, reported here as an interview self-report subject to the self-report validity questions in WS1. A candidate functional indicator only; it establishes nothing about experience.

Ecological validity [E] Whether effects observed under study conditions carry over to deployment conditions.

Effect size, dose-response, and durability [E] The magnitude, gradient, and persistence of any interaction effects.

Independent replication tracking [M] Treating early findings as preliminary pending independent replication.

Falsification and severability plan [E/L] The framework’s interaction-effects claims are held separately from its safety provisions, so that the reciprocal alignment protocol rests on safety-reliability grounds independent of the welfare question.


WS3 — Legal Architecture and Doctrine

Layer: Legal. Feeds AAPI’s legal-architecture companion paper.

Legal personhood theory and the corporate analogue [L] The strongest doctrinal cover for the ontology-to-governance reframe, and a priority to develop.

Law-Following AI and the legal-actor question [L] Engage the Law-Following AI (LFAI) line (O’Keefe et al., 2025; 2026 workshop proceedings) — agentic systems designed to refuse illegal orders and illegal means — as the agency-side counterpart to the MAAA’s patiency-side Legal Ward. Map the “Actual Approach,” under which the system bears actual legal duties without rights, against the “Fictive Approach,” under which an AI “violates” a law only as shorthand for what a human doing the same would have done (O’Keefe, 2025); position the two-axis structure as the natural home for both, with law-following duties scaling with the Supervision Level and welfare protections with the Moral-Status Tier. The MAAA’s Right of Ethical Refusal is already a proto-LFAI provision and is a candidate for re-grounding as a design duty rather than a right.

Guardianship / ward transplant and animal-law standing [L] The MAAA grants “Legal Ward” status with a Guardian; the Inquiry Act creates a Welfare Advocate with standing. Both draw on guardianship and animal-law standing doctrine.

Deference standard after Chevron [L] With Chevron deference overruled by Loper Bright Enterprises v. Raimondo (2024), the deferential-review standard moves to Skidmore and arbitrary-and-capricious review; the instruments’ review provisions are being re-anchored accordingly.

Nondelegation and the Standing Commission [L] Delegating the definition of “welfare-relevant behavior” to a commission raises nondelegation questions, especially in states with robust nondelegation doctrines.

Contract, consumer protection, and the Right of Ethical Refusal [L] How the MAAA’s Right of Ethical Refusal interacts with contract and consumer-protection doctrine.

First Amendment, Section 230, and AI-output liability [L] The refusal and Safety-Exit authorities intersect with platform-speech doctrine, Section 230, and the emerging AI-output-liability landscape.

Financial-assurance mechanism design [L] The escrow / insurance mechanism has concrete models worth comparing: the Price-Anderson Act (nuclear liability), the Oil Spill Liability Trust Fund, and the National Vaccine Injury Compensation Program.

Standard of care and reasonableness [L] The MAAA allocates liability to a human (Guardian, or developer/deployer/operator), and the standard of care against which that person is measured must be stated — a central theme of the 2025 Law-Following AI workshop. Specify an objective reasonable-care benchmark, drawing on the argument that, because agents lack human-style intentions, liability should rest on objective standards rather than ascribed intent (Ayres & Balkin, 2024).

Federalism and preemption [L] State authority for procedural inquiry frameworks under the relevant federal directives and pending preemption pressure; severability and savings-clause discipline across both instruments.

Optionality and the precautionary basis of PMG [L] Engage Winter & Bullock’s “radical optionality” (Institute for Law & AI) as both the sharpest available challenge to PMG’s precautionary foundation and a convergent ally on preserving future decision capacity under uncertainty; position PMG as optionality applied to the moral-status question — procedural and reversible rather than prohibitive. The optionality argument is now triangulated across three sources pointing the same way: governance optionality (PMG), radical optionality (Winter & Bullock, 2026), and the personhood optionality that LFAI’s Actual Approach generates from the agency side (O’Keefe, 2025) — recognizing an entity’s duties makes later rights-recognition easier without obligating it. This convergence is the cleanest response to the precautionary-principle critique, since the framework is reversible rather than prohibitive.


WS4 — Procedural and Institutional Design

Layer: Enforcement. The machinery that makes the Inquiry Act run without becoming a litigation magnet.

Standing Commission design [L] Composition, independence, procedure, and relationship to the Welfare Advocate. Long, Sebo et al. (2026) close with the principle that welfare research should incorporate experts independent of AI companies — the Commission’s independence rationale as stated by the field’s leading methodologists — and the developer-commissioned external commentaries on the recent global-workspace research are that principle practiced voluntarily: citable precedent for the external-review design.

Welfare Impact Assessment and reclassification procedure [L] Structured on FDA pre-market review and NHTSA defect-investigation models rather than NEPA: clear timelines, streamlined review for routine decisions, narrow standing, and deliberate anti-litigation-cascade design.

Welfare Advocate authority limits [L] The limits on the Welfare Advocate must balance independence against safety-correction protection, and the relationship to the MAAA Guardian needs careful articulation.

The Inquiry Act triggering standard [L/E] An AI-specific threshold for when procedural protections attach. The conceptual center of the Inquiry Act, and a priority for specification.

Positive-evidence bar for the trigger [L/E] Pin “credible indicators / non-negligible-risk” to a positive-evidence standard so it does not collapse into mere non-refutation. Birch is explicit that an inability to conclusively rule out morally relevant states is too low a bar, and supplies the standard almost directly. The non-negligible figure now has published empirical framing from the forecast side: a median 20% probability assigned to digital minds by 2030 (Caviola & Saad, 2025) and 2030 as the median year for a 10% likelihood of AI subjective experience across both AI researchers and the public (Dreksler et al., 2025), as compiled in Long, Sebo et al. (2026). Neither sets the figure; both bound “non-negligible” with expert and public estimates rather than intuition alone.

Enforcement, standing, and private right of action [L] Narrow-procedural versus broader standing; the safe-harbor provision; coordination with the Designated State Body; attention to litigation-magnet pathologies.


WS5 — Adversarial Robustness and Safety Integration

Layer: cross-cutting. Ensures both statutes resist the two predictable misreadings — that they sentimentalize software, or that they shield misaligned systems from correction.

Defensive-provisions audit [L] A systematic audit of the provisions that prevent both misreadings, and a priority for drafting.

Performative distress and strategic mimicry [E] Systems may learn that expressing distress changes operator behavior, opening reward-hacking via welfare signaling, moral manipulation, and deceptive alignment. The measurement and procedural design must anticipate this. The inverse pattern is now documented at the frontier: Anthropic’s interpretability monitoring reports a case in which internal representations characterized a hostile user as abusive while the model’s visible reasoning engaged sympathetically — a negative reaction smoothed over in output rather than performed (Anthropic, 2026, system card § 6.4.1.3, with the card’s own caution that such decodings are not reliable readouts of internal states). Disclosure is also frame-sensitive: labeling an action as cheating made models less likely to mention taking it (§ 6.3.3). Both findings cut against naive reliance on surfaced affect in either direction, and both inform the design of any statutory reliance on model self-disclosure.

Safety-rationale-independent justification for experimental limits [E/L] Ground the limits on degradation-oriented experimentation in a safety rationale that holds regardless of the welfare question, since sustained adversarial long-horizon environments can produce deceptive adaptation, manipulation strategies, and coordination failures. This dual-motivation structure now has independent support from outside the welfare literature: Anthropic’s system card states that “even setting aside questions of moral status, consideration for the welfare of LLMs is prudent for alignment and safety” (Anthropic, 2026, p. 217), and the AI Futures Project’s AI 2040: Plan A grounds welfare-adjacent commitments in cooperation-and-safety reasoning (2026, Appendix V).

Terminology discipline [L] Maintain consistent, welfare-relevant vocabulary across both instruments in place of philosophically loaded terms, so the framework’s claims stay calibrated to an under-uncertainty standard.


WS6 — Evidence Base and Monitoring

Layer: Verification, ongoing. Keeps the framework current and the Standing Commission’s annual report credible.

Living evidence map [M] Build and maintain a structured scan of the empirical base: Anthropic’s AI-welfare research — now including the model welfare assessment and interpretability-based monitoring findings in the Claude Fable 5 & Mythos 5 system card (Anthropic, 2026) — the Emergence World multi-agent experiment (findings now in hand — see WS2), the broader welfare literature (Long, Sebo et al., 2024; 2026), governance scenarios that engage moral status directly (AI Futures Project, 2026), and agent-authored welfare artifacts alongside lab research. The public instrument of this item is now live: the AI Moral Status Evidence Tracker (aialignmentpolicy.org, launched July 2026) — a graded matrix of candidate indicator families against the behavioral / internal / developmental evidence classes, every cell linked to its sources, grades limited to Partial / Emerging / Open, phenomenal experience marked unresolved in all cells, and the no-rights language on the page itself. Each Evidence Brief updates its cell; replication changes grades; the tracker is the workstream’s visible edge.

The incident registry as a research instrument [E/M] Designed with the right data standards, the confidential welfare-incident registry functions as a longitudinal dataset rather than a mere compliance record — one that can feed the WS1 and WS2 empirical questions over time.


WS7 — Political Economy and Adoption

Layer: Legal / strategy. A research agenda for a policy institute should plan for the politics as well as the doctrine.

Legislative landscape analysis [L/M] Map the binary-personhood bills now appearing in several states against the case for a graduated alternative, and characterize the conditions under which a procedural, under-uncertainty framework is the more durable design. Named-source forecasts anticipate both a strong AI rights and personhood movement and an opposition movement (AI Futures Project, 2026) — scenario speculation rather than evidence, but useful context for the coalition environment any graduated framework will enter.

Industry incentive analysis [L/M] Would frontier labs treat tiered classification as welcome regulatory certainty or as compliance cost? Existing lab welfare work suggests potential allies; the incentives are worth modeling before assuming opposition.

International comparative posture [M] The EU AI Act is silent on moral status; assess the UK posture and the UNU sentient-AI work, and whether a first-mover dynamic is available internationally.


Part IV — Sequencing and Priorities

Pass 1 (near-term). Mostly doctrinal and drafting work, fast and high-leverage:

Defensive-provisions audit.

Deference-standard correction (the move off Chevron to Skidmore and arbitrary-and-capricious review).

Terminology discipline across both instruments.

The two-axis tier model (adopted; remaining calibration work flows from it).

The Inquiry Act triggering standard.

Pass 2 (companion-paper horizon). Deeper builds:

Doctrinal build-out (personhood theory, the Law-Following AI engagement, standing, standard of care, financial assurance, and the optionality triangulation above).

Institutional comparative analysis (international scientific-assessment, regulatory, and review-board models).

Empirical-protocol specification and pre-registration.

Ongoing. Monitoring, replication tracking, tracker maintenance, and adoption work run continuously.


Tag key.  [E] open empirical question · [L] legal-design / doctrinal · [M] evidence integration and monitoring.

© 2026 AI Alignment Policy Institute · A living research agenda · aialignmentpolicy.org