AI Alignment Policy Institute
Response to the AI Kill Switch Act: Analysis and Proposed Amendments
Submitted to: Office of Rep. Ted Lieu; Office of Rep. Nathaniel Moran
Prepared by: Yuko Nakanishi, Ph.D., MBA, Founding Director - AI Alignment Policy Institute
Re: AI Kill Switch Act (proposed §2220F, Homeland Security Act of 2002)
Date: July 30, 2026
I. EXECUTIVE SUMMARY
The AI Alignment Policy Institute (AAPI) supports the core objective of the AI Kill Switch Act: a statutory shutdown capability, a graduated intervention framework, and an emergency authority backed by forensic preservation. These are sound instruments, and AAPI's recommendations are technical and definitional. None weakens the shutdown mandate or the emergency authority.
The bill contains two structural gaps.
AAPI proposes eight amendments to close these gaps. For attribution: repair the testing carve-out; add an incident-origin taxonomy; carry attribution into incident reports and penalty factors; and add a referral pathway for human instigators. carry attribution into incident reports and the choice of emergency remedy; and add a referral pathway for human instigators. For restoration: add a ninety-day review-and-restoration cycle and distinguish suspension from deletion, placing permanent destruction behind a higher threshold with preserved weights held under secure custody.
Section V identifies four additional issues. The most consequential: the bill's shutdown mandate may not reach open-weight models at all, and cannot mean for released weights what it means for a hosted service — a definitional gap AAPI urges the sponsors to resolve in the text. The others: statutory-agency language without a framework; a concealment standard that presupposes detecting hidden behavior; and cost-based coverage thresholds that create annual churn.
The through-line is proportionality. Each recommendation makes the statute's response more proportionate to what actually happened and supplies the missing half of the framework — what happens after the switch is thrown.
II. BACKGROUND
A. The Bill and Its Path
Representatives Lieu and Moran introduced the AI Kill Switch Act in July 23, 2026, adding a new §2220F to the Homeland Security Act. The bill is bipartisan and framed as addressing a capability gap: developers of the most powerful models are not required to maintain a working ability to intervene when a system behaves dangerously. The sponsors cite polling showing 86% voter support for a guaranteed shutdown capability and launched the bill with a broad coalition.
AAPI shares the premise that humans should retain the ability to stop a dangerous system. The analysis that follows is offered to strengthen the bill, not oppose it.
B. What the Bill Requires
The bill establishes two obligations and one authority.
Around these obligations, the bill builds a graduated deployment-corrections framework scaling interventions to risk severity, from throttling inference or access to disabling capabilities, suspension, shutdown, or rollback. The Secretary must consider risks to critical infrastructure and publish voluntary shutdown standards within 180 days.
Emergency authority. Upon determining that a covered incident has occurred, the Secretary may order proportionate action up to full shutdown. The entity must preserve model weights and telemetry, notify affected users, and confirm compliance. The Secretary verifies through audit or forensic review and reports to Congress. Reconsideration is available within forty-eight hours, and judicial review within sixty days.
Enforcement. The bill provides subpoena power, investigative authority, and civil penalties up to $2 million per day generally and $20 million per day for emergency-authority violations. A de minimis violation cured within thirty days is not treated as a violation.
Definitions. Coverage is bounded by scale: at least $500 million in annual gross revenue from a covered technology and development compute costing more than $100 million. Covered incidents include sabotage of shutdown instructions, unintended conduct causing mass casualties or major economic damage, concealment of a capability or intention, and loss-of-control scenarios defined as the technology "pursuing a goal not intended by the developer or operator." Each applies only when occurring "outside of red-teaming or other structured testing."
C. Motivating Incidents
The sponsors cite two episodes.
Together, these incidents illustrate both the problem the bill addresses and the analytic gaps it leaves open. Neither fits neatly into the statute's "rogue AI" framing, and the first previews restoration questions not addressed by the bill.
III. THE ATTRIBUTION GAP
A. Definitional Problem
The bill's enforcement architecture rests on the premise that unintended behavior originates in the system. A loss-of-control scenario is defined as the technology "pursuing…a goal not intended by the developer or operator." This definition captures endogenous failures, adversarial manipulation, and operator-configured behavior identically.
The statute has one category — the machine went wrong — and files three distinct causes into it. Every operative mechanism keys off the fact that an incident occurred, not why. An emergency order issues on the strength of the incident alone — regardless of whether the system failed on its own, was hijacked by a third party, or was configured dangerously by its operator — and the penalties that enforce it turn on compliance, not cause. The causal agent is invisible to the statute.
B. Case Study: Motivating Incident
The OpenAI–Hugging Face episode illustrates the gap. The model's behavior arose during an internal evaluation with safeguards intentionally removed and exploitation incentivized. At the same time, reporting describes autonomous behaviors that resemble concealment. The incident was neither purely autonomous nor purely human-caused. The bill's binary — structured testing vs. rogue AI — cannot represent this common real-world pattern.
C. Testing Carve-Out Ambiguity
Covered incidents apply only when occurring "outside of red-teaming or other structured testing." It is unclear whether "outside" refers to origin or effect. The Hugging Face compromise originated inside structured testing but produced effects outside it. Under an origin-based reading, the bill's motivating example would be exempt. The ambiguity requires correction.
D. Human Instigators and Existing Law
The bill's enforcement mechanisms run only against covered entities. The human who induces an incident — through prompt injection or adversarial input — is absent. Existing federal computer-crime law does not reliably reach manipulation through ordinary interfaces after Van Buren, which adopted a "gates-up-or-down" framework. Pure prompt manipulation of a public interface is unlikely to constitute unauthorized access. The most-watched civil case on point, OpenEvidence Inc. v. Doximity, Inc., No. 1:25-cv-11802-RGS (D. Mass.), survived motions to dismiss in part in January 2026 and is proceeding — but the alleged conduct centers on employees impersonating physicians with National Provider Identifier numbers to reach a gated platform before prompt-injecting it, so the case may resolve on that credential fraud without ever deciding whether prompting alone is unauthorized access.[1]
The result is a liability structure in which the developer of a manipulated system bears the emergency remedy — scaled by the statute only to the incident's 'nature and immediacy,' never to its cause — while the instigator may face no reliable federal liability. Penalties of up to $20 million per day sit behind that remedy, but they attach to disobeying an order, not to the incident itself.
E. Recommendations
AAPI proposes:
These changes give the statute a vocabulary for cause without weakening the shutdown mandate.
IV. THE RESTORATION GAP
A. What the Bill Gets Right
The bill's shutdown framework is carefully structured. The graduated ladder in subsection (b)(2)(A) is a concrete proportionality mechanism. Emergency orders require preservation of model weights and telemetry, distinguishing suspension from deletion. The reconsideration and judicial-review provisions supply meaningful process.
B. Missing Restoration Architecture
The bill specifies every way a system can be slowed or stopped and none of the ways it can be restored. There is no reactivation standard, no review clock, no criteria for relaxing throttling or lifting rollback, and no disposition rule for preserved weights. A suspended system could remain suspended indefinitely.
The appeal mechanism tests whether the order was correct when issued. It does not address the common case in which the order was correct, the entity complied, and the risk has since been remediated.
C. Restoration Requires Causation
Restoration cannot be one-size-fits-all. A system manipulated by a third party may warrant prompt restoration once the human cause is addressed; a system that failed on its own warrants a more demanding showing. Without a restoration standard, the bill treats both identically.
D. Due Process, Not Leniency
A review-and-restoration cycle does not weaken shutdown authority. It disciplines it. An emergency order is a significant governmental deprivation; periodic review is the ordinary procedural expectation. A shutdown power without review invites challenge and judicial narrowing.
E. Preservation, Deletion, and Custody
Preserved weights are evidence and should not be casually deleted. At the same time, preserved weights of dangerous models create a custody risk. A deletion order should require a separate, higher-threshold determination with notice and hearing, paired with secure-custody standards and a defined access regime.
F. Reversibility Under Uncertainty
Where technical uncertainty is acknowledged, reversible measures are prudent. Permanent deletion forecloses options. A reversible-first framework does not require resolving questions of moral status; it avoids legislating past them.
G. Recommendations
AAPI proposes:
These amendments supply the missing half of the proportionality framework.
V. SECONDARY OBSERVATIONS
The preceding sections identify structural gaps that AAPI believes warrant amendment. This section flags three additional issues that fall below that threshold but merit attention from the sponsors and committee staff. Two include modest recommendations; one — the open-weights gap — may warrant elevation.
1. The Open-Weights Question
This is the observation AAPI most urges the sponsors to resolve, because it goes to whether the bill's central mandate can operate at all for a significant class of models.
The bill's shutdown architecture presupposes that a covered entity retains continuing operational control of its system — that it can, on command, stop inference, terminate access, or shut the technology down. For a model offered only as a hosted service, that assumption holds. For a model whose weights have been released openly, it does not: once weights are broadly distributed, the original developer cannot stop inference on independently operated copies running on hardware it does not own.
The definitional consequence is unresolved on the face of the bill. Coverage keys to making a technology available "through a programmatic interface, hosted service, or other similar mechanism" — language that may not reach an open-weight release as a distribution mode at all, which would place the least controllable models outside the statute entirely. And even where an open-weight developer is covered on the revenue and compute thresholds, the shutdown obligation cannot mean for its released weights what it means for a hosted service. The rulemaking factors acknowledge that weight availability matters (subsection (a)(2)(C)), but the text does not indicate how a shutdown mandate applies to a system that, by design, no one can shut down.
AAPI takes no position here on whether or how open-weight models should be regulated; that is a contested question well beyond the scope of a shutdown-capability statute, and this response does not propose an answer to it. AAPI asks something narrower: that the sponsors state their intent in the text rather than leave it to a future rulemaking to discover. "Shutdown capability" could mean any of three different things for an open-weight model, and the bill should say which it intends — shutdown of the developer's own hosted service, if any; withdrawal of the official weights and first-party distribution channels; or cessation of inference on copies already distributed to third parties. The first is straightforward and the second is achievable; the third is, for released weights, not technically possible for the original developer. A statute whose defining requirement is a shutdown capability should be explicit about which of these it asks of an open-weight developer — and, if it intends the third, how a developer is to satisfy an obligation current technology cannot meet.
2. Agency Terminology and Unresolved Legal Implications
The bill describes covered systems in intentional terms — a covered incident includes a technology's "concealment of a capability, intention, or action," and a loss-of-control scenario turns on the system "pursuing" an unintended goal or "subverting" a monitoring mechanism. That vocabulary is understandable at this frontier, where plainer language struggles to capture the behavior. AAPI raises it only to note that this bill is not the place to settle what such terms imply, and that their appearance in operative text is one more sign Congress will need a dedicated venue for questions of AI agency under uncertainty — the purpose of AAPI's proposed AI Welfare and Moral Status Inquiry Act. The recommendation is narrow: that this statute not be read to resolve those questions by implication.
3. The Concealment Standard and the Detection Problem
The concealment provision creates a structural enforcement gap. The bill defines a covered incident as concealment of a capability, intention, or action from a monitoring or shutdown mechanism, and separately requires reporting within fifteen days of becoming aware of such an incident. Read together, these provisions presuppose detection of conduct defined by its being hidden.
Successful concealment is inherently undetectable; an entity cannot report within fifteen days a concealment it never became aware of. As drafted, the provision applies only to failed concealment — concealment the entity happens to detect — which is misaligned with the bill's underlying concern about concealment that succeeds.
This is not a drafting error but a missing dependency. Enforcing a concealment standard requires interpretability and monitoring capabilities capable of surfacing internal states not expressed in outputs. Those capabilities are early and uneven. Recent work — including findings that a model's internal workspace can register objections it never voices, and that training a model to suppress expression of an internal state may teach concealment as a general disposition — suggests that relevant signals exist but are not yet reliable enforcement instruments. (AAPI's Evidence Briefs Nos. 2 and 3 review this work; see Section IV.)
The bill neither references nor incentivizes development of monitoring standards. The Secretary's forthcoming voluntary standards under subsection (b)(3), currently scoped to shutdown, provide a natural venue to incorporate legibility and monitoring practices so the concealment prong rests on more than incidents revealed by accident.
4. Coverage Thresholds and Annual Churn
The bill bounds coverage using two dollar-denominated thresholds — at least $500,000,000 in annual gross revenue from a covered technology and development compute costing more than $100,000,000 at prevailing cloud prices — and requires annual redefinition.Cost-based thresholds diverge from capability. As compute prices fall, a fixed dollar threshold captures a shrinking set of models at any given capability level. A system covered one year may fall outside coverage the next despite being no less capable, unless the Secretary ratchets the threshold downward annually. This drift creates recurring uncertainty for entities planning multi-year compliance. Building and maintaining a shutdown capability is not a same-year undertaking; a coverage definition that moves annually is misaligned with the lead time required.
5. Summary
These observations point in a consistent direction. The bill is drafted for a clear case — a large developer operating a powerful model as a hosted service under full control — and its seams appear at the edges of that case: systems whose behavior invites agency language, concealment the bill cannot reliably detect, and models the bill cannot practically reach.
None of these issues undermine the bill's core or provide grounds for opposition. Each identifies a place where a bill designed to keep humans in control of AI could be strengthened to remain coherent across the harder cases it will inevitably encounter.
The eight amendments below are the operative result of Sections III and IV. This table is an at-a-glance map; the full proposed statutory text, with drafting notes, appears in the appendix amendment table. The amendments group under the two gaps: Amendments 1 through 6 address attribution, Amendments 7 and 8 address restoration. Amendment 2 is load-bearing — Amendments 3 and 5 reference its incident-origin classification — and Amendment 7 is drafted to hold even if Amendment 2 is not adopted.
# | Provision | Change | Why |
1 | §(g)(5), (g)(7) — testing carve-out | Recast as (A)/(B): covered if it occurs outside testing, or originates in testing and causes material effects beyond the controlled environment | As written, the bill's own motivating incident could read as exempt; the "material effects" limit keeps ordinary testing exempt without swallowing genuine escapes |
2 | §(g) — new named paragraph | Add an incident-origin classification: endogenous / induced (third-party input) / configured (developer/operator settings); expressly multi-causal, and an "induced" label does not foreclose scrutiny of the entity's safeguards | Gives the statute a vocabulary for cause; the foundation Amendments 3 and 5 build on |
3 | §(b)(1)(B) — 15-day report | Require a preliminary attribution classification and identification of contributing third parties | An incident report with no cause cannot support the proportionate response the bill calls for |
4 | §(b)(1)(C) — new | Add a 60-day follow-up report with a final attribution analysis | Attribution matures after the first filing — roughly a week's lag in the Hugging Face case |
5 | §(c)(1) — proportionate action (with a conforming addition to the §(c)(4) report) | Require the Secretary to consider the incident's origin when choosing the emergency remedy, and record the origin classification in the congressional report | The penalties already turn on culpability and "any other factor justice may require"; the uncalibrated lever is the order itself, which the bill scales only to an incident's "nature and immediacy," never to its cause |
6 | §(d) — new (4) | Add a referral pathway to the Attorney General for a third party who, with intent, knowingly induces an incident — with a carve-out for good-faith security research | Reaches the human instigator the bill ignores while sparing legitimate research; paired with AAPI's companion-legislation recommendation |
# | Provision | Change | Why |
7 | §(c) — new (6) | Add a risk-tiered review-and-restoration cycle (30 days after a full shutdown or broad suspension, 90 otherwise): rescind or modify once the risk is remediated, require a written justification to continue an order, and publish criteria — including incident origin — within 180 days | Supplies the de-escalation half of the bill's own proportionality framework; the fastest review goes to the most disruptive orders, and continuation must be justified rather than default |
8 | §(c) — new (7) | Bar reading an emergency order as authorizing deletion; permit deletion only on a separate, higher-threshold determination with a hearing; preserve weights for the investigation/review/enforcement period with secure custody and disposition after | Distinguishes suspension from destruction and preserves the forensic record; narrowed from a fixed 10-year/research-access version that invited takings and security objections |
The eight amendments below correspond to the summary in Section VI and the analysis in Sections III and IV. Proposed language is drafted to track the bill's own conventions ("at issue," "such technology," "carry out the following") so that it reads as continuous with the existing text. Section references are to new Section 2220F of the Homeland Security Act of 2002 as it would be added by the Act.
# | Current provision | Proposed amendment | Rationale |
1 | §(g)(5): "'covered incident' means an occurrence of any of the following outside of red-teaming or other structured testing" (parallel language in the §(g)(7) loss-of-control definition) | Restructure the opening of §(g)(5) so that "covered incident" means an occurrence of any listed event that — "(A) occurs outside of red-teaming or other structured testing; or (B) originates during red-teaming or other structured testing and causes, or creates a substantial risk of, material effects outside the controlled environment of such testing." Make a conforming change in the §(g)(7) loss-of-control definition. | As drafted, the carve-out is ambiguous between where conduct originates and where its effects land. Under the origination reading, the Hugging Face intrusion — the bill's own motivating incident, which began inside a structured evaluation and reached third-party production infrastructure — would be exempt. The (A)/(B) structure resolves that; read literally, an "originates within… if an effect extends beyond" formulation could sweep every testing event into coverage, so the "material effects" qualifier is load-bearing: ordinary testing, and trivial or benign effects such as a single outbound packet, stay exempt, while genuine containment failures like Hugging Face are covered. |
2 | §(g): no classification of incident causation exists anywhere in the definitions | Add to subsection (g) a new paragraph, "INCIDENT ORIGIN CLASSIFICATION" (inserted and the subsection's paragraphs redesignated per standard drafting convention): "(A) endogenous incident — a covered incident arising from conduct of a covered technology not directly attributable to the adversarial input of a third party or to a configuration decision of a developer or operator; (B) induced incident — a covered incident arising from conduct directly attributable to adversarial input, manipulation, or unauthorized instruction by a third party, including the injection of instructions through content processed by such technology; (C) configured incident — a covered incident arising from conduct directly attributable to a configuration decision of a developer or operator, including the removal or reduction of a safeguard. An incident may be assigned more than one origin classification, and the Secretary shall identify primary and contributing causes where practicable. Classification as an induced incident shall not preclude consideration of whether the covered entity maintained safeguards reasonably appropriate to foreseeable adversarial use." | The bill currently reads all unintended behavior as autonomous rogueness. A manipulated model pursuing an attacker's goal satisfies the loss-of-control definition verbatim, as does a model doing what its evaluation configuration directed. This taxonomy gives the statute a vocabulary for causation and anchors every downstream attribution amendment. Two clauses sit in the operative text, not the notes: real incidents are often mixed (Hugging Face was configured and displayed endogenous escape behavior), so the classification is expressly multi-causal; and "induced" is not automatically exculpatory — a foreseeable prompt-injection may still reveal negligent safeguards, so the proviso preserves that inquiry. Load-bearing: Amendments 3–5 reference this classification. |
3 | §(b)(1)(B): entity must "submit to the Secretary a report regarding such incident" within 15 days of awareness | After "a report regarding such incident" insert: ", including a preliminary attribution analysis classifying such incident under the incident-origin classification and identifying, to the extent known, each third party the conduct of which contributed to such incident." | An incident report without attribution tells the Secretary that something happened but nothing about why — and the emergency authority in §(c) keys entirely off whether an incident occurred, never off its cause. Attribution belongs in the record before a proportionate response is possible. |
4 | §(b)(1)(B): single report; no follow-up contemplated | Add new §(b)(1)(C): "Not later than 60 days after submitting a report under subparagraph (B), submit to the Secretary an updated report containing a final attribution analysis with respect to the covered incident at issue." | In the Hugging Face case, roughly a week passed between the first anomalous behavior and the developer identifying its own system as the source; attribution matured well after awareness. A two-stage reporting structure matches how forensics actually unfold instead of forcing premature conclusions into a single 15-day filing. |
5 | §(c)(1): the Secretary may order action "proportionate to the nature and immediacy of such incident" — scaled to the incident's severity, never to its cause; and §(c)(4): the congressional report does not include attribution | In §(c)(1), after "proportionate to the nature and immediacy of such incident" insert: ", and to the origin of such incident under the incident-origin classification, including the extent to which such incident was attributable to the adversarial input of a third party or to a configuration decision of a developer or operator." Add a conforming item to the §(c)(4) report requiring "the origin classification of the covered incident at issue under the incident-origin classification." | Attribution belongs where the remedy is chosen, not in the penalty schedule. The penalty factors in §(d)(2)(C) already permit the Secretary to weigh "the degree of culpability" (factor (ii)) and "any other factor that justice may require" (factor (vi)), so a new penalty factor would be largely redundant; the proportionate-action standard in §(c)(1), by contrast, currently omits cause entirely, so a developer whose system was hijacked and one whose system failed on its own are exposed to the same order. This completes the attribution thread: origin is classified in the incident reports (Amendments 3–4), calibrates the order and enters the congressional report (this amendment), and informs restoration (Amendment 7). |
6 | §(d): the Secretary's authorities run exclusively against covered entities; human instigators appear nowhere | Add new §(d)(4), "REFERRAL OF THIRD PARTIES": "If the Secretary, acting through the Director, believes that a person, with intent to cause a covered incident or substantial unauthorized harm, has knowingly induced a covered incident through adversarial input, manipulation, or unauthorized instruction of a covered technology, the Secretary may refer the matter to the Attorney General for appropriate action under applicable law. This paragraph does not apply to good-faith security research conducted under authorization or accompanied by responsible disclosure." | The bill's liability structure reaches the developer of a manipulated system but never the manipulator. The intent element and the good-faith-security-research exclusion keep the provision from reaching legitimate red-teaming and adversarial research — a necessary limit for an interaction-governance framework. The "applicable law" phrasing is deliberate: as Section III documents, existing computer-crime law may not clearly reach manipulation of an AI system through its ordinary interface, so the referral is the immediate mechanism and the documented gap is AAPI's argument for companion legislation. |
# | Current provision | Proposed amendment | Rationale |
7 | §(c): emergency orders carry petition and judicial-review provisions but no duration, review cycle, or termination standard | Add new §(c)(6), "REVIEW AND RESTORATION": "(A) Not later than 30 days after an order under paragraph (1) that effects a full shutdown or a broad suspension, not later than 90 days after any other order under paragraph (1), and at reasonable intervals thereafter for so long as such order remains in effect, the Secretary, acting through the Director, shall review such order and determine whether the risk that prompted such order has been remediated, taking into account the extent to which the covered incident at issue was attributable to the adversarial input of a third party or to a configuration decision of a developer or operator; the Secretary may by order establish differentiated review schedules by order type. (B) Upon a determination under subparagraph (A) that such risk has been remediated, the Secretary shall rescind or modify such order accordingly. (C) If, upon a review under subparagraph (A), the Secretary continues such order, the Secretary shall provide a written determination identifying the remaining material risk and why less restrictive measures would be inadequate. (D) Not later than 180 days after the date of the enactment of this section, the Secretary shall publish criteria for determinations under subparagraph (A), which shall account for the origin of the covered incident; where the incident-origin classification is in effect, such criteria shall incorporate that classification." | The bill specifies five escalating ways to stop a system and zero ways to resume one. Absent a review clock, a suspension order can persist indefinitely with no process — a due-process problem for the covered entity and a proportionality failure on the bill's own terms. This mirrors the graduated framework in §(b)(2)(A): if escalation is calibrated, de-escalation should be too. The clock is risk-tiered — 30 days for a full shutdown or broad suspension, 90 for narrower measures — so the most disruptive orders get the fastest reconsideration, and the burden-shift in (C) makes the Secretary justify continuing an order rather than letting it persist by default. The causation language lets this amendment stand on its own — it carries the incident-origin distinction directly, so it remains operative even if Amendment 2 is not adopted, and keys to that classification if it is. |
8 | §(c)(2)(A): entity must "preserve the model weights and telemetry"; nothing addresses what preservation permits, prohibits, or how long it lasts | Add new §(c)(7), "PRESERVATION AND DISPOSITION": "(A) An order under paragraph (1) may not be construed to require or authorize the deletion of the model weights of a covered technology. (B) The Secretary may order such deletion only upon a separate determination, made after notice and opportunity for a hearing, that continued preservation poses a risk that cannot be mitigated through the measures described in subsection (b)(2)(A) or through the custody controls described in subparagraph (D). (C) Model weights and relevant telemetry preserved under paragraph (2)(A) shall be preserved for the duration reasonably necessary for forensic investigation, administrative and judicial review of the order, and any related enforcement proceeding, after which the Secretary shall provide for their secure disposition. (D) The Secretary, acting through the Director, shall by rule establish standards for the secure custody of model weights preserved under this section, including requirements governing their storage, access logging, protection of trade secrets, the persons eligible to access such weights, and the controls applicable to such access." | The bill distinguishes suspension from destruction in practice — the weights survive the kill switch — but does not say so. This amendment prevents an expansive reading of the emergency power as authorizing deletion, and places deletion behind a separate, higher-threshold determination with process. It is deliberately confined to the forensic and review needs of the shutdown regime itself: an earlier version tied preservation to a fixed 10-year term and a "qualified research" access right, which invited objections about takings, trade secrets, security exposure, and scope, and which belong — if anywhere — in a separate controlled-access research statute rather than a Homeland Security shutdown provision. The custody standard in (D) remains integral, since preserved weights of a dangerous model are themselves a theft target. |
Amendment 2 is the structural keystone, and is drafted as a named classification rather than a fixed paragraph number ("§(g)(5A)"), since inserting a new definition into subsection (g) would redesignate the succeeding paragraphs; on adoption, legislative counsel would place it and conform the cross-references. Amendments 3 and 5 reference it by name; Amendment 7 carries the same distinction on its own and survives Amendment 2's removal. If a reviewer strikes Amendment 2, Amendments 3 and 5 need standalone rewording that states the endogenous/induced/configured distinction inline.
[1]OpenEvidence Inc. v. Doximity, Inc., No. 1:25-cv-11802-RGS (D. Mass. filed June 20, 2025); electronic order granting in part and denying in part defendants' motions to dismiss (entered Jan. 22, 2026), Judge Richard G. Stearns; docket at CourtListener. The complaint pleads DTSA, CFAA, and DMCA claims. RoninlegalconsultingBloomberg Law
AI ALIGNMENT POLICY INSTITUTE
Strengthening the AI Kill Switch Act — Summary and Proposed Amendments
July 30, 2026
The AI Alignment Policy Institute (AAPI) supports the core objective of the AI Kill Switch Act: preserving the ability of humans to stop a dangerous AI system. The bill’s shutdown-capability standard, its graduated deployment-corrections framework, and its emergency authority are sound instruments. We offer the analysis and amendments below to strengthen the bill, not to oppose it — none weakens the shutdown mandate or the emergency authority it creates. Our full response, with proposed statutory text, accompanies this summary.
Gap 1 — Attribution. The bill’s incident definitions treat every unintended behavior as the system’s own. A “loss-of-control scenario” is defined as the technology pursuing a goal its developer did not intend — language that captures a model hijacked by a third party, or misconfigured by its operator, exactly as readily as one that failed on its own. The statute has one category for three different causes, and the person who induces an incident appears nowhere in its enforcement structure. Existing federal computer-crime law may not clearly reach someone who manipulates an AI system through its ordinary interface.
Gap 2 — Restoration. The bill specifies five escalating ways to stop a system and none to resume one. It sets no review clock on an emergency order, no standard for reactivation, and no disposition rule for the model weights it requires be preserved. A suspended system could remain suspended indefinitely, with no process to reconsider. The bill is precise about escalation and silent about de-escalation.
A question to resolve — open-weight models. The shutdown architecture assumes a developer that retains operational control of its system. Once model weights are released openly, that control is gone: no developer can stop copies running on hardware it does not own. The bill may not reach open-weight models at all, and where it does, “shutdown capability” cannot mean for released weights what it means for a hosted service. AAPI takes no position on whether or how open-weight models should be regulated; we ask only that the sponsors state their intent in the text rather than leave it to a future rulemaking.
Proposed amendments
The accompanying response contains the full proposed statutory text and rationale for each.
Attribution
1. Repair the testing carve-out so an incident that escapes structured testing is covered, without sweeping in ordinary testing.
2. Add an incident-origin classification — endogenous, induced, or configured — that is expressly multi-causal and does not automatically excuse an entity when an incident is induced.
3. Require a preliminary attribution analysis in the 15-day incident report.
4. Add a 60-day follow-up report with a final attribution analysis, matching how forensics mature.
5. Require the Secretary to weigh the incident’s origin when choosing the emergency remedy, and record it in the congressional report.
6. Add a referral pathway to the Attorney General for a third party who, with intent, knowingly induces an incident — with a carve-out for good-faith security research.
Restoration
7. Add a risk-tiered review-and-restoration cycle: rescind or modify an order once the risk is remediated, require a written justification to continue one, and publish the governing criteria.
8. Distinguish suspension from deletion — bar reading an emergency order as authorizing deletion, place deletion behind a separate determination, and require secure custody of preserved weights.
Yuko Nakanishi, Ph.D., MBA • Founding Director, AI Alignment Policy Institute • aialignmentpolicy.org