Skip to content
← Back to blog
Product strategy

How to evaluate mobile app and game ideas before writing the MVP

10 min read

Written by: Jonathan Reis on

A large category is evidence that people use a product, not evidence that a new solo product can win. This is the process we used to reclassify app and game ideas after comparing them with real competitors.

OpenGraph preview image for this article. How to evaluate mobile app and game ideas before writing the MVP

The first ranking of our mobile product ideas looked reasonable. It was also too optimistic. A reminder app sounded simple, a finance app belonged to a category with good subscription benchmarks, and a puzzle game seemed cheap to prototype. Those statements were true and still not enough to justify building any of them.

The correction was to stop ranking categories and start evaluating products. We compared store listings, official competitor pages, reviews, monetization, acquisition channels, operational responsibilities, and our own willingness to work in each domain. Several ideas moved down immediately. A few remained worth testing, but none earned the label “safe bet.”

Start with the user’s actual substitute

The first question is not “who are my competitors?” It is “what would this person use instead of my product today?” The answer is often not an app from the same category.

For an app that remembers birthdays and suggests presents, the substitutes are Google Calendar, contacts, Amazon, Pinterest, Etsy, and wish-list services. Google Calendar has more than 10 billion downloads and recurring events, reminders, widgets, shared calendars, and Gmail integration. A dedicated birthday app is not competing against a small reminder utility. It is asking users to replace a tool that is already installed and trusted.

The same pattern appears in productivity. Google Tasks has more than 10 million downloads and integrates with Gmail and Calendar. Todoist also has more than 10 million downloads, with natural-language entry, habits, calendars, teams, and integrations. A new generic task app does not start at zero; it starts behind products with years of accumulated behavior.

This does not mean established products eliminate every opportunity. It means the new product needs a narrower job. “Remember a birthday” is weak. “Help a small company approve and send gifts to clients without losing the history” is a different product with a different buyer and distribution path.

Separate category demand from available opportunity

A category can be attractive and unavailable at the same time.

Finance is a good example. Monefy has more than 10 million downloads and roughly 194,000 reviews. It already offers quick expense entry, budgets, recurring items, multiple currencies, synchronization, widgets, export, and passcode protection. That validates the need for simple financial tracking. It also makes “a cleaner expense tracker” a weak strategy.

The opportunity has to move from category to workflow. A product for one profession might solve a concrete decision that Monefy does not: how much a beauty professional actually keeps after a booked service, materials, commissions, and taxes. Even then, the category benchmark is only a reason to interview users. It is not proof that the niche will pay.

The same correction applies to health and fitness. Google Fit has more than 100 million downloads. MyFitnessPal has more than 100 million downloads and a food database with over 20 million items. Nike Training Club has more than 10 million downloads and a large library of coached programs. Health and Fitness can monetize well, but a generic workout log also inherits the data, content, integration, and trust burden of the leaders.

Treat competitor reviews as product research

Store reviews are not a complete market study, but they reveal where a product creates friction for real users. Look for repeated complaints, not isolated anger.

The large puzzle games made this visible. Sudoku.com has more than 100 million downloads and roughly 2.46 million reviews. Its listing includes daily challenges, events, tournaments, hints, notes, auto-check, and more than 10,000 puzzles. Reviews repeatedly mention ads interrupting concentration, rewarded ads failing, crashes, and lost progress.

That evidence does not prove that a new Sudoku will win. It suggests a possible positioning: a calm, reliable, offline Sudoku with less interruption and trustworthy progress recovery. The product still needs distribution. A better experience hidden in a store listing with no traffic is not a business.

The same distinction matters for a “puzzle without energy and pop-ups.” Reviews of Candy Crush Saga and Homescapes mention pop-ups, failed ad rewards, misleading advertising, energy friction, and perceived randomness. This is useful evidence for a positioning hypothesis. It is not evidence that players will download a new game or pay for fairness. The first prototype still has to create a second play session.

Check whether the business requires an operation you do not want

The best market opportunity for a team can be the wrong project for an individual developer. We added a personal-fit decision to the ranking because interest changes the probability of finishing and maintaining a product.

Several ideas were rejected even though their markets were plausible. A device-cost calculator would require tariffs, units, appliance specifications, and regional assumptions. A warranty tracker would invite ongoing work around products, manufacturers, documents, and support flows. A MEI obligations app would require learning Brazilian tax and business rules and keeping them current. Those are not merely implementation details. They define the business’s maintenance burden.

Personal fit also has a geographic dimension. If the goal is worldwide exposure, a Brazil-only domain may be strategically wrong even when local demand exists. This does not make the Brazilian market uninteresting. It makes it incompatible with the current distribution goal.

The decision is allowed to be personal. “I do not want to learn this domain” is a valid reason to stop. It should be recorded separately from “the market is bad,” because the two conclusions lead to different future decisions.

Use real scale without worshipping it

Download counts answer one question: has this product reached users? They do not answer whether a solo developer can acquire the next thousand users.

For daily games, NYT Games has more than 10 million downloads and combines Wordle, Crossword, Connections, Spelling Bee, Sudoku, Strands, and other games with streaks, rankings, and a paid archive. Everyday Puzzles has more than 5 million downloads, over 76,000 reviews, daily puzzles, XP, badges, VIP, and a position among top-grossing trivia apps at the time of the audit.

The lesson is not “build a collection.” It is that a daily-game collection is a content operation. It needs editorial review, localization, scheduled releases, retention systems, support, and a distribution brand. Our global version of the idea starts with one procedural daily game, in English and Portuguese, on the web. Sudoku can be the first title because it already exists. Nonogram can be the second. New games enter only after the previous ones demonstrate return and sharing.

This is a much smaller and more testable plan than launching a collection with five modes.

Rank evidence, not excitement

The final table uses separate dimensions:

  • Simplicity: how difficult the first useful version is to build.
  • Specific potential: how plausible the monetization is for the proposed product, after competition is considered.
  • Confidence: how much evidence supports the hypothesis.
  • Risk: technical, operational, legal, content, and distribution risk.
  • Personal decision: whether the developer actually wants to work on it.

The score is useful for sorting, but it is not a probability. A game can score well because the prototype is cheap while having low confidence in acquisition. An app can score well in a Business category while having no defined buyer. Keeping confidence visible prevents those numbers from looking more precise than they are.

Build the smallest evidence-producing test

The next artifact should answer one question.

For a web utility, that can be a landing page and one working flow. For a game, it should be one mechanic with enough polish to produce a second session. For an existing app such as Daily Sudoku, the test may be acquisition and retention rather than more gameplay.

The minimum test we use is:

  1. Write a specific promise for one audience.
  2. Show a price or monetization choice.
  3. Get qualified visitors, not only friends.
  4. Record the core action.
  5. Measure return, sharing, or payment.
  6. Decide whether to continue before adding systems.

For a game, “I like the idea” is weak evidence. The stronger signal is a player starting another round without being asked. For a utility app, a stronger signal is a user bringing real data into the product or attempting to pay.

What the revised process changed

The original shortlist contained ideas with attractive scores but unclear distribution: generic finance, travel planning, reminders, training logs, cozy games, and several classic puzzle clones. Competitive research moved them down because the user already had a capable substitute or because the business required more content and maintenance than the MVP description admitted.

The remaining candidates are not guaranteed winners. They are simply clearer experiments. A global daily-game collection has a defined expansion path. Nonogram can reuse puzzle infrastructure but still faces a leader with 50 million downloads and roughly 897,000 reviews. A strategy game can be prototyped with cards before a map and multiplayer. A lottery tool can be tested on the web, but its basic function is already a commodity.

That is a healthier backlog. It says what to test, what not to build yet, and why a project was rejected. It does not pretend that a spreadsheet of stars can replace users.

Practical checklist

  • Identify the user’s current substitute, including built-in products and manual workarounds.
  • Record at least three direct or indirect competitors.
  • Capture downloads, reviews, pricing, update cadence, and monetization when available.
  • Read repeated negative reviews and translate them into testable hypotheses.
  • Separate category demand from a specific product advantage.
  • Identify the ongoing data, content, legal, and support obligations.
  • Confirm that the developer wants to learn and maintain the domain.
  • Pick one distribution channel before building a large feature set.
  • Test a landing page, prototype, return session, or payment.
  • Keep research snapshots dated and link them to individual product fiches.

The useful conclusion is not that one idea is safe. It is that the next build should be chosen by evidence and personal fit together. A market can be real, monetizable, and still be the wrong project to make.

Related postRebuilding the Daily Sudoku landing page after the Play Store launch9 min readWritten by: Jonathan Reis on