Ratings and feedback
The acquisition half of Youper's revenue gap
In short
Youper had almost no five-star ratings. The CPO asked for a store prompt, and I proposed using the same prompt to ask unhappy users what was wrong. Five-star ratings went from two in February to 166 in total by the end of May. The 254 written answers from the first three months became one of the signals behind the team’s next project.
My role
Lifting the ratings was the CPO's initiative. Collecting feedback in the same prompt was my proposal.
- Ran the benchmark and brought three flows, recommending the most complete
- Classified the 254 written answers, first by hand, then with GPT-4
- Designed, handed off and tested each version, and set up the tracking. The rules were worked out with the CPO and CEO.
Overview
This was one half of a bigger project. Youper wasn’t financially sustainable, and the team worked on revenue per download from both sides: making more from each download, and paying less to get each one. Onboarding and paywalls (opens in a new tab) covers the first half. This case covers the second.
The work ran across three versions in seven months: a store prompt in February, a star question in March, and new rules at the end of August. Five months passed between the second and third while the product was rebuilt.
We started as simple as it could be, and each later version improved how we asked for what we needed.
The problem
Youper had almost none of the five-star reviews it needed. Ratings decide how an app ranks in search and how it looks to someone deciding whether to install, so they were a direct lever on what each new user cost.
Phase one: Ship an incomplete version to get a baseline
What we shipped
We had a few days left in a sprint and had to decide how to use them. Version 11.07 shipped in February with the native App Store prompt. It fired after someone closed the wrap-up of their first mood check-in, and only when that check-in was positive.
We didn’t want to interrupt someone who had just written about their anxiety to ask for a review. So anyone whose entry was a negative emotion never saw the prompt.
Everyone else did, whatever they thought of Youper. There was no internal question, nothing to catch an unhappy answer, and no feedback collected. I recommended the most complete of the three flows, and agreed we could start with the fastest. It gave engineering something to build that sprint, and it gave us a baseline.
What happened
11.07 changed nothing we could see. One-star ratings kept arriving at about twice the rate of five-star ones, the same as before the prompt existed, until late March. The trade bought us a month and a starting line, and the fix was already scoped.
Phase two: Ask inside the app first, and keep the bad answers
What we shipped
The fix followed the benchmark: ask a star question inside the app, and send only five-star answers on to the store. I also proposed that if we were going to interrupt someone to ask how it was going, we should keep the answer when it wasn’t good.
So one to four stars opened a text field asking how we could improve. It was open text, not options, because we didn’t know yet what the answers would be. The Send button stayed disabled until the user typed something, and there was no way out.
The positive check-in gate stayed. It was about empathy: we didn’t want to ask anyone for feedback right after they’d written about a bad day.
What happened
11.08 shipped in March, and the ratings climbed. More useful, complaints started arriving as writing we could read, not as public reviews from people who had already left.
Phase three: 254 answers, sorted by hand, then classified at volume
What I did
In its first three months, the prompt collected 254 written submissions. That’s the set I classified. I sorted the first batch by hand to find the patterns, then turned them into instructions for GPT-4 to classify the rest. We used AI wherever we could at Youper, so it was a natural step. Nine categories came out of it.
What we found
Two categories tied for first at 23.2%, and they meant opposite things. AI performance was real product feedback. It became one of the signals behind the team’s next initiative, alongside our data reviews, anonymized conversations and our own testing.
Early experience wasn’t feedback at all. It was people saying they hadn’t used the app enough to have an opinion, which meant we were asking too soon. We’d suspected that when we shipped, but decided to move forward anyway. Now we had the evidence to justify a fix.
Phase four: Give users a way out, and a reason to ask again
What we shipped
The fix took five months. Between March and August, the team rebuilt Youper around open conversation with an AI, and I was helping design the chat input. I argued for a skip button from day one. The rebuild was the priority, and the call wasn’t mine.
Version 12.03 shipped in August with a skip button, a check for whether someone had already rated or written, a new ask after every five conversations, and tracking for conversation count and subscription status.
What happened
In design terms, the change was one button. Most of the work was in the rules: who sees the prompt, in what state, and what happens after each answer.
The tracking went in at the same time. Every answer arrived with its star score, subscription status and number of conversations. That let us ask subscribers what was missing, and ask everyone else why they hadn’t subscribed.
We didn’t classify a second batch. But after 12.03, we noticed fewer placeholder answers, like a few dots or a single letter typed just to get past the field, “no,” or “idk.”
Impact
Five-star ratings went from two in February to 166 in total by the end of May.
The 254 classified answers gave the team a ranked list of what users found wrong with the product. And by August, the prompt only asked once someone had used the app, never after a negative check-in, and came back after five more conversations instead of treating one refusal as final.
The rating counts come from Youper’s store dashboard, captured at the time. The category percentages come from the 254 answers we classified. I no longer have access to the analytics, so nothing after August 2023 can be retrieved.
Reflections
Ship the rough version, then fast-follow the fix. I suspected the timing was wrong before we had the data to prove it, and unhappy users sat at a form with no exit for five months. Shipping the rough version quickly was still the right call. What I’d change is pushing for the way out as a fast follow, instead of waiting for three months of data to justify it.
Keep the answer when it’s not good. Asking for feedback wasn’t in the brief. But we were already interrupting people to ask how it was going, and the app had no other way for anyone to tell us anything. So I argued we should keep the bad answers instead of throwing them away. It fed the team’s next project, and it gave me the confidence to speak up more in the room.
What I’d look at next: the one-star rise in June. The chart runs to early July, and the last weeks aren’t all good news. Five-star ratings came off their April peak, and one-star ratings roughly tripled through June. Those didn’t come through our prompt, since only five-star answers went to the store, so people were going there on their own. Why more of them did it by June is the first thing I’d look into.





