Accuracy
Every number this app gives you is an estimate, and the estimates carry real error. Here is how much, measured against food that was weighed on a scale — including the result that is worst, and the parts that have not been measured properly yet.
Last updated 12 September 2026
1. The short version
Use it for trends, not for single meals. The error on any one estimate is large enough that a day’s total should be read as “about this much”; the average over a fortnight is worth considerably more than any day inside it, and the app’s adaptive target is built on exactly that assumption.
If you need a number you can rely on for a single meal — because you are counting carbohydrate for insulin, or managing a condition where the exact figure matters — weigh the food and use the packet. That is true of every app in this category. This one says so.
2. Photographing a plate: measured
This is the hardest thing the app does and the one with the published measurement. On 24 August 2026, 30 plates from Nutrition5k — a research dataset in which every dish was weighed on a scale, so what follows is error against ground truth rather than disagreement between models — were each estimated three times, in four configurations.
Mean absolute percentage error
That is a large number and it is the honest one. A 600 kcal plate might be read as 400 or as 900.
The error is not random, which is the useful part
Every configuration tested compressed toward a typical meal: small plates were over-read by roughly 70–100%, large plates under-read by around 50%, and the estimates correlated with the truth at only about r = 0.4. The model has a strong prior about what a plate of food contains and the photograph moves it less than it should.
Two things follow. Photo logging is at its worst on unusually small and unusually large portions, which is worth knowing when you use it. And the fix is a better prompt rather than a better model — switching between the largest and smallest models moved the error by a few points and did not touch the compression.
3. Describing a meal in words: not measured the same way
Text logging is the majority of what the app does, and it has no equivalent weighed-ground-truth measurement. That is a real gap and it is stated rather than filled with an encouraging guess.
What can be said about it: the dominant error is almost certainly portion size rather than food identification. “Two eggs” is unambiguous and the composition data behind it is good; “some pasta” is a range of maybe three to one, and no amount of model quality closes that, because the information is not in the sentence. The more specific you are — weights, counts, the brand — the narrower the error gets, and that is entirely within your control in a way it is not for a photograph.
4. Barcodes and library recipes: not estimates at all
5. Why an inaccurate number is still worth having
Because the target moves with it. The app does not trust a formula to tell you what you burn — it works that out from what you logged against what the scale did, over a fortnight. A systematic bias in your logging partly cancels: if the app reads your meals 20% high, it also computes your maintenance 20% high, and the deficit between them is closer to right than either figure is.
What that does not survive is bias that changes. Logging carefully on weekdays and casually at weekends, or photographing the small meals and describing the large ones, breaks the cancellation. Consistency is worth more here than care.
6. What this page will say next
Known gaps, listed so their absence is not mistaken for a result:
- No weighed-ground-truth measurement of text logging, which is the highest-volume path in the product.
- The photo figure is 30 plates from one dataset, which is enough to rank models and not enough to be a confident population error rate.
- No measurement of how much a more specific description actually narrows the error, which is the single most actionable thing a user could be told.
This page is updated when the numbers are, in either direction. It is not a medical device and none of the above is clinical guidance — see clause 2 of the terms.