We scored 40 dating app openers through FlirtGym's own rubric, twice each, in August 2026. Openers pairing a light compliment with a specific hook scored highest at 4.4/10. "Hey" scored 1.0 — the floor — in every run. The full ranking, method and limitations are below. This measures what our scoring model rates highly. It is not reply-rate data.
What did the study find?
Eight opener patterns, five openers each, every one scored twice through FlirtGym's `rate-my-text` rubric. Ranked by mean opener score out of 10:
| Opener pattern | Mean /10 | Median | Range | SD | Verdicts |
|---|---|---|---|---|---|
| Compliment + specific hook | 4.4 | 3.0 | 3–7 | 1.84 | 3 SOLID, 1 AVERAGE, 6 WEAK |
| Specific profile detail | 3.1 | 2.0 | 2–7 | 2.08 | 2 SOLID, 1 AVERAGE, 7 WEAK |
| Playful mislabel / tease | 3.0 | 2.0 | 2–6 | 1.63 | 2 AVERAGE, 8 WEAK |
| Direct opinion | 2.5 | 2.0 | 2–4 | 0.71 | 1 AVERAGE, 9 WEAK |
| Generic question | 1.8 | 2.0 | 1–2 | 0.42 | 10 WEAK |
| Copied pickup line | 1.4 | 1.0 | 1–2 | 0.52 | 10 WEAK |
| Appearance-only compliment | 1.2 | 1.0 | 1–2 | 0.42 | 10 WEAK |
| Generic greeting ("hey") | 1.0 | 1.0 | 1–1 | 0.00 | 10 WEAK |
The single clearest result: a compliment needs somewhere to go
"You're gorgeous" scored 1.2. "That jacket is doing work, where'd you find it" scored 7. Both are compliments. The gap is entirely in what happens next.
The sub-scores explain it. Appearance-only compliments scored 3.1 on tone — the model does not think they are rude — but 1.0 on specificity and 1.0 on momentum. There is nothing to reply to. Compliment-plus-hook scored 5.6 tone, 4.8 specificity, 4.7 momentum, 4.2 intrigue: the same warmth, with a door left open.
This was the largest single-pattern gap in the study: 3.2 points between two forms of the same gesture.
Pickup lines scored below a plain question
Copied pickup lines averaged 1.4, below generic questions at 1.8, and all ten runs graded them WEAK. "Do you have a map, I keep getting lost in your eyes" scored 1.0 — the same as "hey".
The rubric penalises them on specificity (1.0) and tone (2.2). A line that could be sent to anyone scores as though it was, which is the same judgement a person makes when they receive one.
How consistent was the scoring?
38 of 40 openers received an identical score on both runs. Mean spread 0.05 points, maximum spread 1 point.
That is the design working: the rubric uses anchored bands written so that the same message scores the same way every time, at temperature 0.2. It matters here because it means the ranking above is a property of the rubric rather than of sampling noise.
Limitations — read these before quoting the numbers
1. This is not reply-rate data. It measures what FlirtGym's scoring model rates highly. No claim is made that it predicts what anyone replies to. Treat it as a consistent, documented opinion about structure.
2. The absolute scores run harsh, and the ordering is the finding. Openers were scored with no profile attached, because an opener has no conversation yet. The rubric weights "specific to her" heavily and cannot verify specificity without her profile, so it compresses everything downward — even the best pattern averaged 4.4, and 70% of all scores came back WEAK. Read the ranking, not the raw numbers.
3. Length is a partial confound. Word count correlates with score at r = 0.45 across the 40 openers, and the rubric explicitly penalises over-long messages. Length does not explain the ranking though: copied pickup lines averaged 10.6 words and scored 1.4, while compliment-plus-hook averaged 11.6 words and scored 4.4. Similar length, very different scores.
4. Small corpus. Forty openers, eight patterns, one scorer. Directional, not definitive.
Method
- Instrument: FlirtGym's `/api/wingman/rate-my-text` rubric — Claude Haiku 4.5, temperature 0.2, anchored 0–10 bands, scoring tone, specificity, momentum and intrigue.
- Corpus: 8 patterns x 5 openers, written to comparable length so the variable under test is pattern rather than length.
- Runs: 2 per opener = 80 scores. 0 failures.
- Date: August 20, 2026.
- No user data was involved. Every opener was written for this study. FlirtGym stores no user content.
If you want to check it, the pattern is simple enough to reproduce: score any corpus through the same rubric at temperature 0.2 and compare rankings.
What this means if you are writing an opener
Leave her somewhere to go. That single property separated the top and bottom of this study more cleanly than wit, length, or confidence did. A compliment with a question attached outscored the same compliment alone by 3.2 points.
Specificity comes second, and it has to be real: the model rewards openers that reference something only that person's profile could have prompted. Everything at the bottom of the table — "hey", "you're gorgeous", a copied line — fails both tests at once.
More on writing openers · Practice them in FlirtGym
Score your own openers on the App Store
Frequently Asked Questions
Which dating app opener scores highest?
In this study, openers that pair a light compliment with a specific hook scored highest, averaging 4.4 out of 10. Openers referencing a specific detail from her profile came second at 3.1. Generic greetings like 'hey' scored 1.0 out of 10 - the floor - in every single run.
Do pickup lines work?
FlirtGym's scoring model rates them near the bottom: copied pickup lines averaged 1.4 out of 10, barely above 'hey' and below a plain generic question. All ten runs graded them WEAK. The model penalises them for the same reason people do - an obviously copied line signals it was sent to many people.
Are compliments a good opener?
Appearance-only compliments scored 1.2 out of 10, second worst of the eight patterns tested, and were graded WEAK in all ten runs. The same compliment with a specific hook attached scored 4.4 - the highest in the study. The compliment is not the problem; the dead end after it is.
How was this study conducted?
Forty openers across eight patterns were scored twice each through FlirtGym's own rate-my-text rubric (Claude Haiku 4.5, temperature 0.2, anchored score bands), for 80 total scores on August 20, 2026. Full method, corpus and limitations are on this page, and the script is reproducible.
Does this show what women actually reply to?
No, and that distinction matters. This measures what FlirtGym's scoring model rates highly. It is not reply-rate data and no claim is made that it predicts real-world responses. It is useful as a consistent, documented opinion about opener structure, not as evidence about outcomes.
Is the scoring consistent?
Yes, notably so. 38 of the 40 openers received an identical score on both runs, with a mean spread of 0.05 points and a maximum spread of 1. The rubric is built with anchored bands specifically so the same message scores the same way every time, and it held up.
More from FlirtGym
- What FlirtGym is · Pricing · How practice works
- Compare: All AI dating apps · vs. RIZZ · vs. ChatGPT · vs. SwipeMatch AI · vs. reply generators
- Guides: How to text a girl · Opening messages · When she stops replying · Keeping a conversation going · Flirting over text · Dating profile tips · For shy guys · Cold approach
- Research: Opener Score Study — 40 openers, 80 scores
- Profile & matches: Why you get no matches · Hinge prompts · Asking her out