Skip to main content

Author tools · Local workspace

Skill quality workbench

Make instructions clearer, define observable tests, and keep the evidence behind your revisions.

Drafts and results stay in this page until you choose to save or export. This workbench does not send them to SkillDB or a model. Browser saves are unencrypted, shared by anyone using this browser profile, and are not tied to an account. Leaving the page loses unsaved work. Normal site visit analytics may still run.

One browser save slot. JSON includes full draft and result snapshots; add cases and record outcomes before saving. Nothing is saved automatically. Import limit: 2 MB.

1. Draft your skill

0 / 50,000 characters · 0 lines · approximately 0 words

2. Review text checks

Deterministic checks update as you type. They do not run the skill, validate YAML, infer intent, or predict model quality. A clear report does not prove the skill works.

What is checked?

Unclosed fences/front matter; repeated headings; TODO markers; common section names; identical actions using Always/Never or Must/Must not; and exact “Return only JSON” style format declarations. Fenced code, blockquotes, indented code and front matter are excluded from instruction checks. Conditional phrasing and semantic contradictions require manual review. At most 80 findings are shown.

Paste a draft or load the example to begin.

3. Define observable test cases

Use a normal task, a boundary or missing-input case, and an adversarial case. Expected results should be specific enough for another reviewer to judge. This page does not execute prompts.

0 / 30 cases · 0 normal · 0 edge · 0 adversarial

No cases yet. Start with a common task your skill should handle.

4. Record evaluation evidence

Run a case with the current draft in your chosen model or harness. Paste the observed output and judge it against your expectation. Record model/version, tools, system context, temperature and seed where available; note anything unknown. Snapshots help reproduce the setup, but stochastic outputs may differ.

Requires a draft and a case. Confirm the draft above is the one you evaluated. Up to 20 results per workbench; export an archive before starting a new one.

Recorded results (0)

0 result(s) match the current draft exactly; 0 use another draft. Verdicts are self-reported evidence, not a quality score or independent verification.