Six characters. Pick one and start talking.
Each character shows up differently. Pick one and start.
Type anything, or pull your last message to Buddy. The model reads it and scores all seven emotions. Nothing leaves this browser.
Two different measurements, because they answer different questions.
| On 24 real messages like the ones people send here | 42% |
| the previous version scored | 17% |
| On the public tweet benchmark (6,590 unseen) | 68.4% |
| always guessing "joy" would score | 30.3% |
The gap between those two numbers is the point. Public benchmarks are short tweets where people state the feeling β "i feel awful". Real messages describe a situation and leave the feeling implied. A model can look good on one and fail the other.
Where it still fails: it cannot handle a turn. "I like her but she has stopped replying" reads as positive, because it sees the good words and misses that the second half reverses them. It also chases emotion words β the phrase "surprise party" makes it answer surprise when the feeling is joy.
How it improved: not by getting bigger. Pretraining it ourselves made it worse. What worked was 3,532 training examples describing situations without ever naming the feeling.
Model card and full evaluation on Hugging Face β
dair-ai/emotion Β·
GoEmotions