waif — naming the feeling in a piece of text
Text in, a reading out: three axes anchored to published human norms, a feeling named in two stages by a calibrated model, and the four designs tried first.
Overview
Write something — a message you have not sent, a note to yourself — and waif tells you what the feeling in it is, what you are doing with it, and every number behind both.

It runs on Jev, a System One model that returns calibrated probabilities rather than text. The interesting part was never whether a model can read tone. It is that "what emotion is this" has no right answer while the questions underneath it do, and that finding out which of those questions to ask took four designs and a scored probe set.
Four designs, one probe set
The first version named the feeling from a table of thirty coordinates I placed by hand. Replacing hand-placed guesses with published human ratings was the obvious improvement, so it was tried — properly, and scored on the same 24 texts against a list of acceptable words for each one.
| Design | Score |
|---|---|
| Nearest word in the whole vocabulary, by published valence/arousal/dominance | 3/24 |
| Nearest word within a family, by rank on the axis that separates that family | 9/24 |
| Nearest word within a family, by distance in published valence/arousal/dominance | 12/24 |
| Family chosen by the model, and the shade chosen by the model within it | 22/24 |
The norms turned out to be a poor way to pick a word, and the measurement said so before I had a chance to argue. Three things went wrong at once, and all three are worth knowing before reaching for an emotion lexicon:
- Dominance is nearly a copy of valence. Across the emotion words used here it correlates with valence at +0.87, so the third dimension carries far less independent information than the theory promises.
- Negative emotions occupy almost no space. Fear, frustration, worry, terror, jealousy and embarrassment sit inside a ball small enough that the nearest neighbour is close to arbitrary, and the ratings are means from about twenty people with a standard deviation near 1.7 on a 1–9 scale.
- A word rated alone is not the same measurement as writing read in context. People rate the word "gratitude" as far more activated than a grateful message reads. Forcing the two onto one scale imports that mismatch as error.
So the norms do the job they are actually good at: anchoring the rubric. Every level of every axis names words whose ratings were measured — "as activated as rage or panic" is a claim you can check, where "very aroused" is not — and the vocabulary table on the page shows where each of the 62 words sits, with its rater count and spread.
Two questions, asked at the same time
Naming the feeling is two decisions, not one. Which family is it in — anger, fear, sadness, shame? And which shade of that family exactly? Families are genuinely alternatives to one another, and once a family is settled its shades are a short list of near neighbours rather than sixty-two strangers.
Neither is asked second. The shade question needs to know the family, which looks like it needs the first answer in hand before it can be written. It does not: there are eleven families and all eleven shade lists are already in my own source file, so all eleven questions go out in the same request, each stating its own family as a premise — if this is anger, which anger? — and code keeps the answer whose premise held. Ten answers are thrown away.
That trade works because Jev reads the text once and answers every question against it in parallel, so ten wasted questions cost tokens and almost no time, while waiting to learn which one to ask costs a whole round trip. Measured on the probe set: 398ms against 762ms, faster on 46 of 46 head-to-head pairs, same family every time, and the extra tokens come to five hundredths of a cent a reading. It was two sequential requests, for a day.
What the split is actually worth is not what I claimed. The argument for it was that one Choice over sixty-two near-synonyms divides its own vote between wordings, so annoyed / irritated / frustrated would come back split three ways. Measured, it does not: one Choice over all 62 scores 44 against this design's 43 on the same 46 readings, and is the more certain of the two (mean 0.812 against 0.689 on the word it picks). Both designs flatten on the same four texts, because what flattens them is the text sitting between two words, not the length of the list.
The split earns its place on something else: two separate doubts. Is this sadness or affection is a different failure from it is clearly shame, but guilt or embarrassment, and the page says different things in each case. One Choice returns one number that cannot tell them apart. The cost is honest too — a wobble in the family answer corrupts the word, because once the family says sadness, nostalgia is not on the ballot. That cost one text in 23.
And sometimes the answer is two words. A text can genuinely be between guilt and embarrassment; the shade Choice comes back 52% and 41%, and that split is the reading. Two words get named, stronger first, when both conditions hold: the top word did not settle on its own (under 0.60) and the pair holds 80% or more of the weight between them.
Both conditions do work. Drop the first and every clean reading acquires a second word at 3%. Drop the second and a text spread thinly across guilt, regret and shame gets reported as a pair when it is really none of them — that one keeps a single word and says plainly that it did not settle. The thresholds come off the probe set: an unsettled shade puts at least 22% on its runner-up, a settled one at most 19%. The rule fires on 5 texts in 23.
Or more is load-bearing, and it took a bug to find. Every probability Jev returns is a whole hundredth — 6,796 values checked, not one otherwise — so a top-two sum does not approach the threshold, it lands on it. The page's own flat example does: apathy 46% and emptiness 34%, printed as adding to 80, and silently dropped by a rule that read more than. My probe set contained no such case in 46 readings, which is exactly why a quantized scale needs the boundary decided deliberately rather than discovered.
What you are doing, which is not what you are feeling
The first version was missing it. Intent had been scattered across two Nouls, is it aimed at the reader and is it asking for something, which were fragments of a judgment rather than a judgment. Both were replaced by one Choice over eight speech acts: venting, asking for help, seeking reassurance, confronting, thanking, deciding, telling you what happened, thinking it through. These are alternatives to each other, which is exactly why a Choice is right here and wrong for emotion words.
Every other Noul was cut, leaving only the gate that decides whether there is anything to read at all. Is more than one feeling present went on measurement — it came back at or above 0.6 on twenty-one of twenty-four texts, which is a property of writing rather than a signal, and what it reached for is read off the shade Choice's own spread anyway. Is it about something not yet happened and is it being held back went on use: each could only ever append one line, neither changed the word, the sentence or a meter, and a question whose whole effect is an occasional footnote costs a reader more attention than it returns.
The reading is one sentence for the same reason. The axes say how it feels and the intent says what you are doing with it, so the intent finishes the sentence rather than sitting beside it as a second verdict about the same text.
What it refuses to do
It has 62 words and no more, and no coordinate separates feelings that differ only by what caused them. Where the shade lands on two words, it names both rather than picking one and sounding certain; where it lands on none in particular, it names the strongest and says it did not settle. Where even the family is unsettled, it says so. Where an axis clears its confidence threshold but the top two levels are a coin toss, that is caught too — confidence is computed over the whole distribution, so a 50/49 split can look peaked and still be undecided.
It is an instrument, not help. Nothing you type is stored or logged.
What is under it
- Frontend: plain HTML and JavaScript, no framework.
- Model: Jev, one call — six questions measuring the text, and eleven speculative ones naming the shade — through a server endpoint (a Cloudflare Pages Function) so the API key never reaches the browser.
- Data: a 62-word extract of Warriner, Kuperman & Brysbaert (2013), norms of valence, arousal and dominance for 13,915 English lemmas.
- Tests: fifteen, against answer objects recorded from real calls, so the suite needs no key and no network.