An average of the wrong things
Dish quality is one diluted component of a number being used as a dish recommendation.
Case study 05 · Product improvement
People search for biryani. The platform ranks restaurants. That mismatch is small enough to ignore on any single order and expensive enough to matter across millions, because a 4.3★ restaurant can still serve a mediocre version of the one dish you came for. This is a proposal for dish-first discovery: ranking the thing the user actually asked for.
An independent analysis of a public product experience. Not affiliated with or commissioned by Zomato, and no internal data was used. The argument rests on the observable behaviour of the discovery flow.
01 / The idea
A restaurant rating is an average over everything a kitchen makes, plus delivery, packaging and service. It is a genuinely good number for choosing a restaurant, and it is being used to answer a much narrower question: who makes the best version of this one dish? Those two questions have different answers more often than the interface admits. A specialist on 3.9★ may make the definitive biryani while a 4.5★ multi-cuisine kitchen makes a forgettable one, and ranked by restaurant score the specialist loses.
Dish quality is one diluted component of a number being used as a dish recommendation.
A kitchen that does one thing exceptionally cannot out-rank a broad kitchen that does everything adequately. The ranking has no way to see the specialism.
A bad biryani becomes a lower restaurant rating. The system never learns the biryani was the problem, so it recommends it again.
Mapping the current flow shows the failure is not one bad screen. It is a substitution that happens silently at step three and is never corrected afterwards.
These aren't three separate risks. They're one sequence. Mismatch is survivable; mismatch that teaches the user to distrust the ranking is not, because the ranking is the product.
02 / The proposal
Two changes, and the second is what makes the first survive. Rank search results by a dish-level score rather than a restaurant one, and collect feedback at the dish level so the score has something to learn from.
Rank the specific item the user searched for, using signals about that item, so a specialist can out-rank a generalist on the dish they specialise in.
Ask about the dish, not just the order. Without this the ranking has no correction signal and degrades into the same restaurant average it replaced.
03 / Challenges
This is an outside analysis of a public product. I have no access to ranking data, order histories or experiment results, which rules out proving anything with numbers. So the case had to be built from observable behaviour of the discovery flow and from mechanisms that either hold or don't, rather than from a metric I cannot see. It is a weaker kind of evidence and it is the honest one available.
Ranking by dish is easy to say. The score needs dish-level signals, and the current feedback loop only produces restaurant-level ones, so on day one there is nothing to rank with. That is why the feedback change is not a nice-to-have bolted on at the end. Without it the dish score has no input, and the proposal quietly collapses back into the restaurant average it was meant to replace.
Every extra question after an order is friction, and most people will skip it. A design that assumes enthusiastic dish-by-dish reviewing is not a design, it is a wish. The realistic version collects sparse feedback and has to produce a usable ranking anyway, which is a constraint I would rather state than design around.
04 / Limitations
A new dish or a new restaurant has no dish-level signal at all, and ranking them below everything established is its own unfairness.
Most orders will never produce a dish rating, so early scores rest on small samples and are noisy in exactly the way ratings should not be.
"Chicken Biryani", "Hyderabadi Biryani" and "Special Biryani" may or may not be the same thing. Ranking a dish means first deciding what counts as that dish.
Any signal that decides ranking gets manipulated. Dish ratings would be worth manipulating from day one.
No experiment, no uplift number. This is a reasoned proposal, and whether it works is an A/B test I cannot run.
Hygiene, delivery time and packaging do not disappear because the ranking is dish-first. A real design has to carry both.
05 / What I learned
The bug was not a screen, it was a proxy quietly standing in for the real signal. Mapping the journey step by step is what made an invisible swap visible.
I started with "rank by dish" and hit the fact that the data to do it does not exist. Most ranking proposals are really proposals to collect something new.
A score with no correction signal decays. The feedback half is not a phase two, it is the thing that keeps the first half alive.
With no internal data, the temptation is to invent a plausible number. Saying plainly that this is reasoning and not measurement is what keeps the rest of it credible.
The complete write-up carries the journey maps, the ranking-signal breakdown and the full argument.