It was never in the room
Check this first, every time, before reading a word of your own file. An instruction file that was not loaded produces exactly the session you would get from an instruction file that was ignored, and the two are indistinguishable from the outside.
Each tool will tell you, and each tool has a different way of being asked:
| Tool | How to see what loaded |
|---|---|
| Claude Code | /memory lists the instruction files with paths, /context shows them under Memory files |
| Cursor | The rule appears in the context pill on the message it applied to |
| Copilot | The references list on the response names the instruction files used |
The reasons a file goes missing are dull and specific, which is good news: a stray CLAUDE.local.md turning off the team's AGENTS.md, a Cursor rule with neither alwaysApply nor a matching glob, a Copilot applyTo pointing at a directory you renamed. None of these errors. They all just quietly subtract your file from the request.
In our experience most reported cases of a model ignoring instructions end here, at a file that was never read. It is worth ten seconds before it is worth an afternoon.
The rule was not checkable
A model cannot act on a sentence it cannot test its own output against. This is the single most common content failure, and it survives in instruction files because the lines look sensible when you read them back:
| Not checkable | Checkable |
|---|---|
| Write clean, maintainable code | Functions under 50 lines. Extract rather than nest past 3 levels |
| Follow our conventions | Named exports only. No default exports |
| Handle errors properly | Every await on a network call is inside a try/catch that logs the URL and rethrows |
| Be careful with the database | No raw SQL. Use the repository in src/db/, and never deleteMany without a where |
| Write good tests | Each new function gets one test for the failure path, not only the happy one |
The test is mechanical and takes a second per line: could a reviewer hold the diff next to this sentence and say yes or no? If two reasonable reviewers could disagree, the model can disagree too, and it will disagree in whatever direction its training data leans.
The left column is not useless, exactly. It is a preference, and preferences belong in a conversation where you can react. What goes in a file that loads on every request should be the things you would enforce in review. Constraints and specificity are the two Academy lessons that drill this directly.
The codebase disagrees with the file
Your instruction file says named exports only. Four hundred files in the repository use default exports. The model has read some of them, because it was asked to change one, and now it has a rule in the prompt and four hundred counterexamples in the evidence.
It will often follow the code. That is not defiance, it is the same instinct that makes it match your indentation without being told: the surrounding code is the strongest available signal about what this project wants, and usually that heuristic is right.
So an aspirational rule needs to say that it is aspirational, and say what to do about the existing files, which is a thing you know and the model cannot guess:
## Exports
Named exports only in new files. Most of `src/legacy/` still uses default
exports. Do not convert them as a drive-by, and do not copy the pattern into
anything new.Three sentences, and it now beats the counterexamples because it accounts for them. A rule that pretends the legacy directory does not exist is a rule the first honest look at the repository disproves.
It was buried
Instruction files grow by accretion. Every incident adds a line, nothing ever removes one, and eventually the file is eight hundred lines of accumulated grievance sitting at the top of every request.
Attention is finite and it is spent. A rule competes with the other rules in the file, with the code that was just read, and with the actual question. At some length the marginal line is not merely ignored, it is diluting everything above it.
The repair is subtraction, and the order that works is this:
- Delete anything the model already does. A line telling a model to write valid syntax is pure cost.
- Move anything path-specific out of the always-on file, into whatever your tool calls a scoped rule.
- Move anything occasional into a command or prompt file that you invoke when you need it.
- For what remains, cut every sentence explaining why. You are writing instructions, not persuading anyone.
What is left is usually a quarter of the original and works better than the whole did. Front-loading the context that matters is the same discipline applied to a single prompt.
You described the shape, not the output
The last diagnosis is the subtlest. A rule phrased as a prohibition tells the model what not to produce and leaves the replacement unspecified, so it improvises one, and the improvisation is the thing you end up complaining about.
"Do not write long functions" produces the same function with a comment apologising for its length. "Split anything past 50 lines into named helpers in the same file" produces helpers. Same intent, one of them names the artefact.
The general form is an output contract: say what the finished thing looks like rather than which mistakes to avoid. In an instruction file that means naming the format, the file it goes in, and the command that proves it worked.
A ten-minute repair loop
Put together, the whole diagnosis is a short sequence, and it is worth running end to end rather than stopping at the first plausible cause:
- Confirm the file loaded. If it did not, nothing else in this list matters.
- Find the specific rule that was not followed. Not the file, the line.
- Read that line as a reviewer. Could you fail a pull request against it? If not, rewrite it until you could.
- Ask whether the codebase contradicts it. If it does, say so in the rule and name the exception.
- Ask whether it is a prohibition. If it is, restate it as the thing you want produced.
- Then, and only then, consider that the file might be too long for any of it to land.
Almost every step is about the writing rather than the tool, which is the honest summary of this whole blog: these files are prompts with a filename, and they fail for the reasons prompts fail. If you want the rewrite in step three scored before you commit it, that is what the desk is for.