← All posts
§ Blog

Owner, verifier, gate

A generated SKILL.md is not a skill until it has an owner, a verifier, and a gate. Agent-authored files can score worse than no skill at all.


Owner, verifier, gate

I asked a model to write the first skill. It produced a card that would look employed on a marketplace. I threw it away.

That scene is already in The description is the trigger. Li et al. was enough, then, to stop me shipping a model-written playbook for a job I already do on Thursdays. The folder is still a sketch. I am not going to pretend it is live.

The sharper line is not “no average benefit.” It is worse than nothing.

Generated SKILL.md versus a platform object with an owner, a verifier, and a gate

Generated files can hurt

Peng, Zhang, et al. (Aug 2026, arXiv 2608.17587) write it as a capability gap, not a copy-editing miss. Expert-written skills help. Agent-authored skills perform 8 to 11 points worse than using no skill. They cite SkillsBench. Following a skill and improving one from execution are distinct capabilities.

That is the news. The rest of SkillsBench is already on this site. Pointer, then move.

WER, their framework, trains a Skill Optimizer outside a frozen executor. The optimizer proposes a skill. A frozen agent executes it. A programmatic verifier scores the outcome. The score is the environment, not the same model grading itself.

WhatNumberWhose run
Agent-authored skill vs no skill8 to 11 points worseSkillsBench, as cited by WER
WER vs no-skill baseline, BFCL v4 Pass@168.83 to 76.63 (+7.80)Peng, Zhang, et al.
WER vs no-skill baseline, τ²-bench Pass@1+3.85same paper
Same 4B backbone, untrained optimizer vs WER-trained+9.35 BFCL v4, +10.29 τ²-benchsame paper
Trained 4B optimizer on BFCL v476.63%, ahead of the off-the-shelf models they used as optimizerssame paper
10% plausible-but-wrong experience, τ²-bench Pass@182.5 to 77.2Zhu et al. 2026, cited by WER
Self-verification after that poison83.3 to 83.2Zhu et al. 2026, cited by WER

I did not train a 4B optimizer. I did not run BFCL. The table is their receipt, not mine.

The poison row is the one enterprises should keep. A loop that asks the writer to mark its own homework recovers almost none of the loss. Self-check is not a verifier.

A skill is not a markdown file

A SKILL.md is a file. A skill, on a platform, is an object with three properties.

An owner. Someone who will put a name on the description and take the miss when it fires wrong.

A programmatic verifier. Something that scores the outcome in the environment: the terminal state, the database, the required action. Not a paragraph that says “looks good.”

A load gate. What is allowed to execute this file, and when. The description-as-trigger note was the sentence that opens the playbook. This note is the object that is allowed to run at all.

I already drew identity as the gate for a callable tool in Identity is the gate. That was whose token is on the call. This is the artifact. Adjacent. Not the same thesis.

WER’s move is to keep the executor frozen and put the score outside the writer. The optimizer never acts in the tool environment. Its entire output is a document. Whether that document helped is decided by a deterministic check against a reference state. Freezing the executor keeps the writer out of the trajectory. Using a verifier keeps a model out of the score. Neither alone is sufficient. That is their sentence. It is also the HiTL: the verifier and the gate, not a person sitting in the thread.

DimensionGenerated file, unsupervisedPlatform object
OwnerThe model that wrote it, which is nobodyA named person who will initial the miss
VerifierThe same model, or a self-checkThe environment, scored without a second LLM
GateIf the file exists, it can runWhat is allowed to execute, and when
Unattended morningA plausible procedure that can make the job worseA file that may not load, or a miss you can see

A generated SKILL.md without those three is not a late draft. It is an unowned procedure with a path on disk.

Self-check is not a verifier

Inference-time loops look like governance. Draft. Run. Revise. The file gets better, or it looks like it does. The writer does not. The next task starts from the same base model, asked again to turn a log into a procedure.

WER is honest about the limit. Their benches have programmatic verifiers. Open-ended work, unseen tools, a different executor: they do not claim the transfer. That is not a footnote. If you cannot score the outcome without asking a model, you do not yet have their loop. You have a conversation about a file.

The poison result sits next to that. Ten percent plausible-but-wrong experience is enough to drop the score. Asking the system to verify itself barely moves it. So “the agent refined the skill” is not an audit sentence until you can name what scored the revision.

The Thursday folder still does not exist. I will not ship a generated file into a hole I have not owned.

What enterprises could steal

You do not need their 4B optimizer. You need three names on the artifact.

  1. Name the owner. A SKILL.md with no person on it is a generated file. The owner is who initials the miss when it fires wrong, or when it should have fired and did not.
  2. Name the verifier before you name the model. Terminal state, database row, required action. If the only score is another LLM, you are still inside the writer.
  3. Name the gate. What is allowed to execute this file, and when. Existence on disk is not a gate.
  4. Do not ship a generated SKILL.md unsupervised. Let the model draft. Throw the draft away, or put it behind the verifier. The 8 to 11 point drop is the citation for “worse than nothing.”
  5. Do not call self-check a verifier. A loop that recovers a tenth of a point after poisoned experience is not a gate. It is a mirror.

What this post is not

This is not a remake of the trigger note. That note was the sentence at the top of the file. This one is the object that file has to become before it is a skill.

It is not “WER beats GPT.” A trained 4B optimizer ahead of off-the-shelf writers is a number on their table. The distinction is owner, verifier, gate.

It is not a claim that I ran their benches, or that the Thursday folder is live. The overnight artifact is still the next receipt.

The product question is no longer whether you have a SKILL.md. It is whether anyone owns it, what is allowed to execute it, and what scores a revision.


Field note from the build-in-public log. NDA-safe, no client names, rounded figures only. If this matches what you are seeing, get in touch.

§

BELOW THE LINE

PODČÁRNÍK · SIDEWAYS GLANCE, NOT A SUMMARY

On the colleague who scattered SKILL.md every evening. Kids with toys, going mental, right before bed. Then the adult cleans that shit up.

A skill is an SOP. A policy in one file. In a firm it would be a methodology on slides. Here it has a .md suffix and pretends someone already did the work.

A document landed. Title at the top, numbered steps, “verified” at the bottom. In the preview it looked like meeting minutes. The font was a circular that used to go round the enterprise before email. No alarm.

He opened it for two seconds. The bullets were enough. That is how minutes get read. The whole thing only when the house is on fire, and nothing was burning.

He wrote thanks. Short, no question mark. Official thanks. You put that under minutes so it shows they arrived, not that anyone read them.

He forwarded them into the thread where minutes go. The attachment went with them. Unread. For minutes that is not a sin. It is a mess dressed as operations.

In the corridor they asked if it actually works. He said he has it in his mail. He did.

Minutes are not tested. Minutes are sent.

ČESKY — ORIGINAL PODČÁRNÍK

O kolegovi, který každý večer sypal SKILL.md. Podobný bordel, jako tuhle Večerníček uměl udělat s papíry

Skill je SOP. Směrnice v jednom souboru. Ve firmě by to byla metodika na slajdech. Tady to má příponu .md a tváří se, že už to někdo odpracoval.

Přišel mu dokument. Název nahoře, očíslované kroky, dole „verified“. V náhledu přílohy to vypadalo jako zápis ze schůzky. Font jako z oběžníku, který obíhal podnik ještě před mailama. Žádný poplach.

Otevřel to na dvě vteřiny. Stačily odrážky. Zápis se čte takhle. Celý jen když hoří dům, a nehořelo nic.

Napsal díky. Krátké, bez otazníku. Úřední díky. Pod zápis se to píše, aby bylo vidět, že to dorazilo, ne že to někdo četl.

Přeposlal to do vlákna, kam zápisy patří. Příloha šla s ním. Nepřečtená. U zápisu to není prohřešek. Je to bordel, který se tváří jako provoz.

Na chodbě se ho zeptali, jestli to fakt funguje. Řekl, že to má v mailu. Měl.

Zápis se nezkouší. Zápis se posílá.