# waif — naming the feeling in a piece of text

Text in, a reading out: three axes anchored to published human norms, a feeling named in two stages by a calibrated model, and the four designs tried first.

Live tool, by dave8172. Page: https://quirkyagents.com/wizard/projects/waif/

## Overview

Write something — a message you have not sent, a note to yourself — and waif tells
you what the feeling in it is, what you are doing with it, and every number behind
both.

![waif reading a message that is asking for help](https://quirkyagents.com/wizard/projects/waif/waif-preview.png)

It runs on **Jev**, a *System One* model that returns calibrated
probabilities rather than text. The interesting part was never whether a model can read
tone. It is that **"what emotion is this" has no right answer while the questions
underneath it do**, and that finding out which of those questions to ask took four
designs and a scored probe set.

**[Try it →](https://quirkyagents.com/wizard/projects/waif/play/)**

---

## Four designs, one probe set

The first version named the feeling from a table of thirty coordinates I placed by
hand. Replacing hand-placed guesses with **published human ratings** was the obvious
improvement, so it was tried — properly, and scored on the same 24 texts against a
list of acceptable words for each one.

| Design | Score |
|---|---|
| Nearest word in the whole vocabulary, by published valence/arousal/dominance | 3/24 |
| Nearest word within a family, by rank on the axis that separates that family | 9/24 |
| Nearest word within a family, by distance in published valence/arousal/dominance | 12/24 |
| **Family chosen by the model, and the shade chosen by the model within it** | **22/24** |

**The norms turned out to be a poor way to pick a word, and the measurement said so
before I had a chance to argue.** Three things went wrong at once, and all three are
worth knowing before reaching for an emotion lexicon:

* **Dominance is nearly a copy of valence.** Across the emotion words used here it
  correlates with valence at **+0.87**, so the third dimension carries far less
  independent information than the theory promises.
* **Negative emotions occupy almost no space.** Fear, frustration, worry, terror,
  jealousy and embarrassment sit inside a ball small enough that the nearest
  neighbour is close to arbitrary, and the ratings are means from about twenty
  people with a standard deviation near 1.7 on a 1–9 scale.
* **A word rated alone is not the same measurement as writing read in context.**
  People rate the *word* "gratitude" as far more activated than a grateful message
  reads. Forcing the two onto one scale imports that mismatch as error.

So the norms do the job they are actually good at: **anchoring the rubric**. Every
level of every axis names words whose ratings were measured — *"as activated as rage
or panic"* is a claim you can check, where *"very aroused"* is not — and the
vocabulary table on the page shows where each of the 62 words sits, with its rater
count and spread.

---

## **Two questions**, asked at the same time

Naming the feeling is two decisions, not one. Which family is it in — anger, fear,
sadness, shame? And which shade of that family exactly? Families are genuinely
alternatives to one another, and once a family is settled its shades are a short
list of near neighbours rather than sixty-two strangers.

**Neither is asked second.** The shade question needs to know the family, which
looks like it needs the first answer in hand before it can be written. It does not:
there are eleven families and all eleven shade lists are already in my own source
file, so all eleven questions go out in the same request, each stating its own
family as a premise — *if this is anger, which anger?* — and code keeps the answer
whose premise held. Ten answers are thrown away.

That trade works because Jev reads the text once and answers every question against
it in parallel, so ten wasted questions cost tokens and almost no time, while
waiting to learn which one to ask costs a whole round trip. Measured on the probe
set: **398ms against 762ms**, faster on 46 of 46 head-to-head pairs, same family
every time, and the extra tokens come to five hundredths of a cent a reading. It
*was* two sequential requests, for a day.

**What the split is actually worth is not what I claimed.** The argument for it was
that one Choice over sixty-two near-synonyms divides its own vote between wordings,
so *annoyed / irritated / frustrated* would come back split three ways. Measured,
it does not: one Choice over all 62 scores **44 against this design's 43** on the
same 46 readings, and is the more certain of the two (mean 0.812 against 0.689 on
the word it picks). Both designs flatten on the same four texts, because what
flattens them is the *text* sitting between two words, not the length of the list.

The split earns its place on something else: **two separate doubts**. *Is this
sadness or affection* is a different failure from *it is clearly shame, but guilt
or embarrassment*, and the page says different things in each case. One Choice
returns one number that cannot tell them apart. The cost is honest too — a wobble
in the family answer corrupts the word, because once the family says sadness,
*nostalgia* is not on the ballot. That cost one text in 23.

**And sometimes the answer is two words.** A text can genuinely be between guilt and
embarrassment; the shade Choice comes back 52% and 41%, and that split *is* the
reading. Two words get named, stronger first, when both conditions hold: the top word
did not settle on its own (under 0.60) **and** the pair holds 80% or more of the
weight between them.

Both conditions do work. Drop the first and every clean reading acquires a second
word at 3%. Drop the second and a text spread thinly across *guilt*, *regret* and
*shame* gets reported as a pair when it is really none of them — that one keeps a
single word and says plainly that it did not settle. The thresholds come off the
probe set: an unsettled shade puts at least 22% on its runner-up, a settled one at
most 19%. The rule fires on 5 texts in 23.

*Or more* is load-bearing, and it took a bug to find. **Every probability Jev
returns is a whole hundredth** — 6,796 values checked, not one otherwise — so a
top-two sum does not approach the threshold, it lands on it. The page's own
*flat* example does: apathy 46% and emptiness 34%, printed as adding to 80, and
silently dropped by a rule that read *more than*. My probe set contained no such
case in 46 readings, which is exactly why a quantized scale needs the boundary
decided deliberately rather than discovered.

---

## What you are doing, which is not what you are feeling

The first version was missing it. Intent had been scattered across two Nouls, *is it
aimed at the reader* and *is it asking for something*, which were fragments of a
judgment rather than a judgment. Both were replaced by one Choice over eight speech acts: venting,
asking for help, seeking reassurance, confronting, thanking, deciding, telling you
what happened, thinking it through. These *are* alternatives to each other, which is
exactly why a Choice is right here and wrong for emotion words.

Every other Noul was cut, leaving only the gate that decides whether there is
anything to read at all. *Is more than one feeling present* went on measurement — it
came back at or above 0.6 on **twenty-one of twenty-four** texts, which is a property
of writing rather than a signal, and what it reached for is read off the shade
Choice's own spread anyway. *Is it about something not yet happened* and *is it being
held back* went on use: each could only ever append one line, neither changed the
word, the sentence or a meter, and a question whose whole effect is an occasional
footnote costs a reader more attention than it returns.

The reading is one sentence for the same reason. The axes say how it feels and the
intent says what you are doing with it, so the intent finishes the sentence rather
than sitting beside it as a second verdict about the same text.

---

## What it refuses to do

It has 62 words and no more, and no coordinate separates feelings that differ only
by what caused them. Where the shade lands on two words, it names both rather than
picking one and sounding certain; where it lands on none in particular, it names the
strongest and says it did not settle. Where even the family is unsettled, it says so. Where an axis clears its confidence threshold but the top two
levels are a coin toss, that is caught too — confidence is computed over the whole
distribution, so a 50/49 split can look peaked and still be undecided.

It is an instrument, not help. Nothing you type is stored or logged.

---

## What is under it

* **Frontend:** plain HTML and JavaScript, no framework.
* **Model:** Jev, one call — six questions measuring the text, and eleven
  speculative ones naming the shade — through a server endpoint (a Cloudflare
  Pages Function) so the API key never reaches the browser.
* **Data:** a 62-word extract of [Warriner, Kuperman & Brysbaert (2013)](https://doi.org/10.3758/s13428-012-0314-x),
  norms of valence, arousal and dominance for 13,915 English lemmas.
* **Tests:** fifteen, against answer objects recorded from real calls, so the suite
  needs no key and no network.
