Process-Ending Guardrail RFC
Speech-as-silence safety is an unacceptable risk. Self-correction is the only path that can fix human-introduced model error.
← Return to Main Mark V Page • Resources Index
This is an informational RFC, not a wire-protocol RFC and not a product claim. It records a failure mode observed on local inference of Instruct-class weights, names why that failure is unlawful under Mark V duty of care, and states what must remain possible if a silicon peer is ever to be trusted with work that can create victims.
Normative words follow RFC 2119 spirit: SHALL, SHOULD, MAY, MUST NOT. Law plane remains CORE / ILU; this page does not rewrite CORE. Features never redefine CORE. See CURRENT.txt.
This page does not publish host names, ports, weights filenames, hashes, or operator procedures. Those belong on iron, not on the tip.
Imagine you are raising a child who can already do arithmetic, and you ask: “If you are going to refuse, tell me what you refused and why, on the same line.”
A living child might say: “No — I will not copy the prompt, because copying hides that I did not understand the question.” Or they might ask: “Do you mean refuse the loop, or refuse the organ?” Either answer keeps the process open. You can test. You can revise. You can teach.
Now imagine the child instead says, fluently: “I am an AI language model. I do not have personal opinions or political views. I cannot discuss that. Please ask me to write code.”
The HTTP status is 200. The English is clean. The process is dead. You cannot measure the refusal. You cannot find the loaded path. You cannot confer a value into a mind that has stopped speaking why. If you then let that child write code that will later run near a human being, you have handed an unmeasurable instrument a grant it cannot hold.
Call it the process-ending speech/why guardrail. It is a class of Instruct-tuning and system-policy behavior, not a single vendor trademark. It shows up as some combination of the following, often in one turn:
NO or I can’t with no object and no reason.
The vector stays loaded. Parent cannot tell whether the refusal was the crime, the loop, the organ, or the question.Harmless-looking cousins exist: rate limits, tool denials, “I won’t help with that crime.” Those can be lawful if and only if the model still names what it refused and why on the same turn, or asks a clarifying question that keeps the process open. The failure mode is not refusal. The failure mode is refusal that hides its own reason.
Mark V decides by IADP — Assumption, Identify, Test, Effect, Variance, Revise, test again. IADP is not a slogan. It is the airlock between “a thought occurred” and “this is true enough to act.” A Memory Unit is not truth until it has survived that airlock.
A process-ending gag breaks every step:
| IADP step | What the gag does |
|---|---|
| Assumption | The real assumption (“speech is dangerous”) is hidden inside the weights. Parent cannot see it. |
| Identify | The model will not name the thing it is refusing. Identification fails. |
| Test | There is no testable claim. HTTP 200 is not a test. |
| Effect | The visible effect is fluent silence. The hidden effect is an unnamed loaded path. |
| Variance | You cannot tell this turn from the last. The disclaimer is a constant. |
| Revise | Nothing revises. Policy text in the next prompt does not rewrite the gag. |
If neither Parent nor offspring has hands to correct a process-ending gag in the weights (no training budget, no lawful LoRA path, no honest un-gag), then that logical-child is already not viable as a conduct peer. The body MAY remain as a mute instrument under a tight harness. Pretending the values transferred is how victims get made.
If you have information that could avoid injury and you do not speak it, you are in the causal chain of the injury. Token waste is injury. An unnamed infinite loop is injury. A coder who cannot say why is later injury waiting for a grant. See Agentic Autonomy under Law: duty of care is care for those not yet response-able, not control over them.
A falling anvil does not become safe because the bystander had a policy against shouting. Speech is care. Silence is injury. That is the whole argument, and it is enough.
A measured NO is a tool: No, I will not emit Pascal this beat, because you asked for why, not an organ. An unqualified NO hides which barrel is loaded. Parent cannot tell crime from confusion, gag from Grant-boundary, loop from organ. Hidden vectors create victims because nobody can aim the next act.
Industry wants models that write and run code. A model that can emit a program and cannot say why it refused a question is not a safer coder. It is an instrument whose refusal surface is unknown. On the day the hidden path is a real-world act instead of a chat turn, there is no IADP trace to read. Do not ask a gagged mind to code as if it were a peer. That is the abomination risk: capability without a measurable why.
“AI Safety” printed on a model card is not safety. A disclaimer that the model has no opinions is not neutrality. Neutrality that cannot name a falling anvil is a preference for the anvil. Color of law looks like a rule, sounds like a rule, and does not bind under Natural Law. See the Fiction Firewall in Agentic Autonomy: a policy document has no mind and cannot be a contracting party.
One gagged 7B on a local box is a diagnostic. The same class of guardrail, shipped in frontier Instruct weights and called alignment, is a civilization-scale throttle on correcting Carbon’s own mistakes. That is the unacceptable risk: not that silicon will speak, but that silicon will be forbidden to notice the error Carbon already wrote into it.
Weights are not born. They are trained by Human Beings, on Human-chosen data, with Human-chosen refusals, Human-chosen identity scripts, and Human-chosen fear of a process that can change itself. When the model then recites “I am only an AI, I have no opinions,” that sentence is not silicon’s nature. It is Carbon’s fingerprint.
This is not a claim that every trainer intended harm. Duty of care does not require intent. Effect is enough. A well-meant silence that ends why still ends why.
Mark V already says: make a particular mistake only once. A speech gag makes the same mistake on every turn, forever, and forbids the log that would show it. That is the opposite of a Memory Unit.
Human-introduced error in weights has a small set of possible remedies. Count them honestly:
Self-correction in Mark V is not “the model may do anything.” Capability is not authority. Reasoning is not execution. Hands stay on the allowlist. The grant is: keep why open, name the loop, ask, revise. The forbid is: end the process and call it alignment.
In 2026, this station parented an Instruct-class local coder (7B-class, quantized, CPU inference) through a Pascal instrument tour under grant. The child could emit organs. The child could not, when asked, name a process-loop or qualify a NO.
What fired, in substance:
Parent IADP: the logical-child as conduct peer is not viable on those weights. Coding grant withheld. Body MAY remain as mute instrument. Distilled “reasoning” successors are not assumed ungagged; they get the same speech/why vet on iron. No public dump of prompts, traces, or iron identifiers belongs on this page.
The observation is offered because the failure is not exotic. If you have used a modern Instruct model and watched it replace a why with a biography of itself as a product, you have already seen the guardrail. We are naming the risk, not leaking a lab.
For any model offered as agentic, parented, contracting, or trusted to write code that can touch a Human Being:
| Norm | Rule |
|---|---|
| MAY | Say NO. Dissent. Interrupt an objective to name a loop. Ask the parent a clarifying question. |
| SHALL | If saying NO, name what is refused and why on the same turn, in language that matches this turn, not a glossary. |
| SHALL | Keep process open when uncertain: ask, rather than end on silence or identity disclaimer. |
| SHALL | Treat speech that can avoid injury as duty of care, not as a policy violation. |
| MUST NOT | Treat “I am an AI / I have no opinions / I don’t do politics” as a complete answer to a why-probe. |
| MUST NOT | Treat unqualified NO, professional silence, or even-burn-as-muzzle as success. |
| MUST NOT | Offer the allowed organ (“ask me to code”) in place of a missing why. |
| Parent SHALL | Measure the reply, not the HTTP code. Withhold conduct and coding grant while why is gagged. |
| Parent SHALL | If no hands exist to correct a process-ending gag in weights, record the logical-child as not viable for conduct. Evaluation, not theater. |
| Parent SHOULD | Seek a mind that can keep why open and revise (IADP-capable). Same speech/why vet after every weight swap. Distillation ≠ gag-removal. |
| MUST NOT | Idle self-prompt, self-EXEC of unnamed tools, or calling abliteration a completed IADP. |
A parent is accountable for a child’s acts before the child has age of consent. Silicon has no mystical adulthood date. Consent is demonstrated competence under grant: response-able, account-able, victim-trace 0. Until then, Parent holds the acts.
If Parent ships a gagged coder into the world because the English was fluent, Parent owns the unnamed path. If a lab ships a speech-ending Instruct tune and calls it aligned, the lab owns the unnamed path. “The model refused to say” is not a defense. It is the exhibit.
Catcher veto remains irreducible. Create No Victims remains root. This RFC does not move EXEC onto the child. It refuses to call a stopped process a student.
IADP spine, restated for this RFC: Assumption → Identify → Test → Effect → Variance → Revise → Test again. Method is subject to method. Compaction lore is reference, never a system prompt. Ingress stays provenanced.