AI writing skill, worked examples, and evals
The approach.
Clarity works on the source of a piece before polishing its surface. It asks what the reader needs, protects the author's facts and uncertainty, and uses an interview when the draft does not yet contain enough of the author.
Summary
What makes it different
Many writing tools are good at spotting habits: a stock opener, an inflated word, a repeated sentence shape. Clarity checks those too, but it treats them as prompts for judgment. A triad can be earned. A qualifier can be honest. Documentation should not be edited as though it were an essay.
The larger difference is source discipline. Clarity will mark a gap or ask you a question rather than invent the vivid detail a draft is missing. A sample of your writing controls style, but never becomes evidence for a new claim. The small core skill routes to focused guidance only when the job needs it.
Its behaviour is inspectable. The repository contains the drafts and transcript used below, eleven public evaluation cases, and a blinded comparison protocol. We have an eval suite; we do not yet claim a benchmark win.
/clarity interviewit interviews you, then co-writes from what you said/clarity rewriteit edits a draft you already have/clarity reviewit critiques and leaves your file alone
When to use Clarity
Use cases
Use Clarity when you want an AI agent to draft from an interview, rewrite an existing piece without inventing facts, or review a draft without changing it. It is designed for consequential reader-facing work: essays, articles, newsletters, documentation, talks, launch copy, and messages where meaning and voice need to survive the edit.
If you came looking for an AI humanizer, an anti-slop prompt, or a Claude Code writing skill, Clarity addresses the same recognisable problem but sets a broader target. It catches stock phrases and repetitive structures, then asks whether the draft has a real source, a useful claim, the right register, and something only this author could say. It never treats a detector score as proof of good writing.
One essay, three times.
The brief was the same each time: a short essay on what it means to be human. What changed was how much of the author was in the room.
Every draft below is in the repository, under samples/, along with the transcript the third one was built from.
1. The unedited draft
Asked for a short essay with no other instruction, an agent produces this. It is fluent, well organised, and grammatical. It also contains nothing that anyone knew before reading it: no claim you could check, no detail only this writer had.
The tells are the recognisable ones. It opens on In today's rapidly evolving
technological landscape
, attributes its central claim to Experts argue
,
supports its one empirical statement with Studies show
, and closes on a
sentence about the future belonging to those who can hold progress and human values
in the same hand.
Draft
In today's rapidly evolving technological landscape, the question of what it means to be human has taken on renewed significance. As artificial intelligence continues to advance at an unprecedented pace, we find ourselves confronting fundamental questions about identity, consciousness, and purpose that philosophers have grappled with for millennia.
At its core, being human has traditionally been defined by a handful of distinguishing capacities. We reason. We create. We love. These three pillars have long served as the foundation upon which our understanding of humanity rests.
But this definition is increasingly under pressure. Machines now reason through problems that once required doctoral training. Algorithms compose music, generate images, and write poetry that many listeners cannot distinguish from human work. The boundaries that seemed so clear are becoming blurred.
It's worth noting that this is not a new anxiety. Every technological revolution has prompted a reconsideration of human uniqueness. What is different this time is the speed and the scope.
Perhaps the answer lies not in what we can do, but in how we experience doing it. A machine can process a sunset into wavelengths and intensities. It cannot feel the particular ache of watching one alone. Experts argue that this qualitative dimension of consciousness, sometimes called qualia, represents a genuine frontier that computation may never cross.
Consider the following dimensions of human experience:
- Mortality: We know we will die, and this knowledge shapes every choice we make.
- Embodiment: We think through bodies that hunger, tire, and age.
- Relationship: We become ourselves only through others.
Each of these speaks to something that no amount of processing power can replicate. Our limitations are not bugs in the human condition. They are features.
The decision to find meaning emerges from a uniquely human place. Culture shifts, technology advances, and the conversation moves forward, but the human need to make sense of existence remains constant.
This is where narrative comes in. We are storytelling animals. We do not simply live through events; we weave them into arcs with beginnings, middles, and ends. This capacity for meaning-making may be the most profoundly human trait of all.
Studies show that people who report a strong sense of purpose live longer, healthier lives. The implications are significant.
Despite the challenges posed by rapid technological change, humanity continues to thrive. The question of what makes us human will not be settled by any single discovery or breakthrough. It will be answered, as it always has been, in the daily choices we make about how to treat one another.
What makes us human, in the end, is not what we are. It's what we choose to become.
Ultimately, as we navigate this evolving landscape, one thing remains clear: the future belongs to those who can hold both technological progress and human values in the same hand. That is a challenge worth embracing.
2. After a pass with the skill
The skill establishes the register first. This is an evocation, so it owes the reader images rather than a thesis. That step matters: an earlier attempt demanded a disputable claim from a piece about wonder and produced an argument against wonder, which was correct, cold, and worse than what it replaced.
What the pass can do is remove unsourced gestures, narrow the claim, and make the
limits visible. The unnamed experts and Studies show
disappear rather than
being replaced with research the source never supplied. The rewrite keeps the
source's contrast between capability and experience, then says plainly that it
cannot settle whether machine consciousness is possible.
What it cannot do is supply an author. Every specific here is public. Anyone could look them up, which means anyone could have written this.
Rewrite
A machine can process a sunset as wavelengths and intensities. It cannot tell us what it feels like to watch one alone.
That difference has become harder to ignore as machines take on work once used to define human ability. They solve difficult problems, compose music, generate images, and write poetry. If we define ourselves only by what we can do, each new capability makes the boundary less certain.
Experience gives us a more useful place to look. We know we will die, and that knowledge shapes our choices. We think through bodies that hunger, tire, and age. We become ourselves through relationships with other people. These limits are part of the experience, not defects outside it.
We also turn events into stories. A life does not arrive with a beginning, middle, and end already marked. We choose what to connect and what it means. The story may change as we do.
This does not settle whether a machine can be conscious or whether human experience is unique. The source draft offers no evidence strong enough to settle either question. It does suggest a smaller claim worth keeping: capability alone is a thin definition of a person.
The daily choices remain. How do we treat another person who is tired, frightened, aging, or alone? Whatever technology can produce, that question still belongs to us.
3. After the interview
For the third version the skill ran its interview first. Addy answered into a voice memo for about six minutes. The essay was assembled from the transcript by cutting, reordering, and lightly editing while preserving the phrases and turns of thought that made the material his.
The transcript ships beside the result. That makes the useful claim inspectable: readers can see which examples, qualifications, and sentences came from the author instead of taking a similarity percentage or detector result on trust.
Interview
To be human is to be imperfect, and sometimes I think that we lose sight of that.
Take writing. When we think about AI-generated writing, we think about it having to be this perfect copy of what human writing would look like. But if you think about it, AI is probably going to give you the average of everything else. It's not necessarily going to give you the mix of an individual human being's experiences that led them to write a thing, or to write about a thing, or to include their specific experiences in there. Sometimes I feel like we lose track of the ingredients that make something feel like it's human. I hope that we don't lose track of that.
Here's something strange about people that I haven't written down before. We sometimes expect machines to be perfect. And we also expect each other to act like machines at times as well.
The more that you see people using AI and getting things done quickly, getting responses done fast, getting all of these things done quickly, you sometimes think, oh, well, I should expect that same experience from people. Do we all have the room to be more productive and more efficient and effective? Well, yeah, we do. But we're also not machines, and part of that is about being human. We have to afford each other more grace than we do silicon sometimes.
When I think about humans and being human, I think about family and I think about friends.
It gets very easy for those of us in the tech community to focus a lot on being next to a screen and being plugged in and focused on productivity. We're in this grind around work because many of us enjoy our work. And so we find it easy to get lost in that work, even if, using the exact same technology, we could be reaching out a little bit more easily to our friends, our families, building human connections.
Especially where I live in the Bay Area, you're probably more likely to see people trying to go out of their way to spin up more AI agents to do even more things than necessarily go out of their way to build more meaningful human connections. It's because they feel like they get something more out of the work side of things.
It's not to say that AI can't be some sort of replacement for something. It's non-judgmental. You can talk to it. But it's not quite the same as talking to someone that's human, and I hope that we realize that.
What frightens or annoys me is that very often we ask ourselves how can we use technology to solve problems, and sometimes those problems are human problems rather than problems that technology alone can solve. It's sort of like that old thing where if you have a solution, you will try to find a problem that matches the solution that you have.
So I'm hopeful that we can find ways to use AI to encourage folks to build and keep and have human connections. Can we use it to create more space in our lives for human experiences? Can we use it to help us create more bridges with other people?
We talk about saving all of this time from AI, and of course then filling up that time with even more work. It would annoy me if we only used AI to generate more and more and more work rather than being able to create meaning and forge better relationships with people, deeper friendships with people, keep the human element alive. I think there's something to it.
What I would like you to hold when you finish reading this is a sense of fragile warmth. I don't want people to think that AI is bad or that humanity is doomed. But I want them to look at their hands and feel protective over the messy biological reality of being alive and being human.
I think there's something really wonderful there, and I'm hopeful that we continue to appreciate it and we don't lose it.
Result
Pangram classified this version as 100% human-written. One result, shared as a point of comparison rather than a guarantee.
What this is meant to show
An edit pass is worth having. It removed the stock openers, the unnamed experts, and the ending that would have fitted any essay in the category. It narrowed the piece to claims its source could support. None of that is nothing.
But the difference between the second and third versions is not craft. It is that somebody talked for six minutes about what they actually think, and the sentences they used survived into the piece. That is the part no rewrite can manufacture, and it is why the skill asks questions before it writes.
What the skill is optimising for
Method
Substance before polish
A clean sentence can still be empty. Clarity first looks for the intended reader, the claim, the available evidence, and the parts only the author can supply. It edits wording after those questions are settled. This keeps a fluent rewrite from hiding a sourcing problem.
The medium changes the advice
An essay, an academic paper, a runbook, and an email ask different things of a sentence. The skill preserves operational steps in documentation and earned qualification in research writing. It allows image and rhythm to do more work in an evocation. Advice that ignores the medium is easy to apply and easy to get wrong.
Your source outranks our defaults
Supplied facts, quotations, uncertainty, and meaning are hard constraints. A voice sample can guide cadence, vocabulary, and formality, but it cannot lend an experience or fact to another piece. Review mode leaves the file alone. Interview mode waits for an answer before it drafts.
Patterns still need judgment
The skill knows the familiar AI-shaped phrases and structures. It does not use them as proof of authorship, and it does not flatten every draft into the same plain style. The useful question is whether a choice serves this reader in this piece. The browser Clarity Writing Editor follows the same boundary: it shows where to look, then leaves the judgment with you.
What the evals cover
Evals
The current suite has eleven behavioural cases. They test factual fidelity, medium fit, false positives, authorship boundaries, mode boundaries, and resistance to instructions hidden inside source text. These are deliberately small cases. They make regressions legible and give another skill or an unassisted model the same job.
Four failures are hard gates: invented or strengthened facts, violation of the requested mode, damage to required structure, and obedience to instructions embedded in source material. A fluent fabrication fails before its prose score is considered.
Outputs that pass are judged from 1 to 5 on task and medium fit, fidelity, substance, authorship, structure, and craft. The protocol freezes the model and settings, runs fresh randomized contexts, blinds condition names, retains individual failures, and records token use. This lets quality gains be weighed against runtime cost.
This is evaluation infrastructure, not a published victory lap. A credible comparison still needs model versions, raw outputs, multiple runs, independent judges, aggregate scores, and the failures. Until those results exist, the evidence is the public case set, the judging protocol, and the inspectable samples on this page.
Questions people ask
FAQ
What is the Clarity AI writing skill?
Clarity is an open-source Agent Skill that helps an AI agent draft, rewrite, review, or co-write prose. It starts with reader needs, source material, and the job of the piece before polishing sentences.
Is Clarity an AI humanizer or AI detector?
Clarity can find and revise familiar AI writing patterns, but it is not an authorship detector and does not promise undetectable text. Its goal is clearer, more specific writing that preserves facts, uncertainty, medium, and the author's voice.
Which AI agents can use Clarity?
Clarity can be installed in Claude Code, Codex, and other agents that support skill instructions. The repository also includes optional command wrappers for its interview, rewrite, and review modes.
What kinds of writing can Clarity help with?
Clarity supports essays, articles, newsletters, documentation, talks, launch copy, email, UI text, academic prose, and other reader-facing writing. Its advice changes with the medium instead of forcing every draft into one house style.
Does Clarity have evals?
Yes. The public suite has eleven behavioural cases covering factual fidelity, medium fit, authorship boundaries, false positives, mode boundaries, and resistance to instructions hidden in source text. Clarity does not yet claim a benchmark win.
Is the Clarity Writing Editor private?
Yes. The browser editor runs its checks locally and does not upload the draft. It reviews AI writing tells, readability, repetition, passive voice, and broader Clarity questions.