Most App Store keyword research goes wrong in the same place. Someone opens a tool, sorts a list by a column called "traffic" or "difficulty", picks the rows near the top, and pastes them into the title. Three weeks later nothing has moved, and nobody can say why — because nothing in that process recorded what was actually true at the time.

The fix is not a better score. It is knowing, for every term you act on, which part of what you are looking at was *observed* and which part was *calculated*.

Observed and modeled are not the same kind of fact

An observation is something a store told you at a moment in time. "On 14 September, in the US storefront, searching `habit tracker` returned your app at position 23." That is checkable. It has a store, a country, a language, a search depth and a timestamp attached, and if you repeat it tomorrow you may get a different answer — which is itself information.

A model is an interpretation layered on top. Traffic scores, difficulty scores, opportunity scores: none of these are published by Apple or Google. Every one is somebody's formula applied to observations. That does not make them useless. It makes them arguments, and an argument is only as good as the evidence under it and the version of the formula that produced it.

The practical rule: **never let a modeled number be the only reason you changed something.** If you cannot state the observation underneath it, you are not doing research, you are reading tea leaves.

Start from the terms your app already touches

Before generating candidate keywords, find the ones where you already have a position. These are the cheapest wins in ASO, because ranking 14th for a term is a fundamentally different problem from ranking nowhere.

Three sources, in order of usefulness:

**Your own listing.** Every noun and verb in your current title, subtitle and description is a term Apple has already indexed you against. Search each one and record where you land. You will usually find two or three terms you rank for that you never deliberately targeted.

**Autocomplete.** When you type a prefix into App Store search, Apple returns completions in an order. That order is not random and it is not alphabetical — it reflects what Apple considers worth suggesting in that storefront. It is one of the few genuinely public signals of relative demand. Record the position, not just the presence: a term suggested first behaves differently from the same term suggested eighth.

**The apps that outrank you.** For any term where you appear below position 10, look at the titles of the apps above you. Their titles are a list of terms someone with more traction decided were worth the characters. That is competitive intelligence you get for free.

Judge a term on three things, in this order

**1. Is it the same intent as your app?** A term can have real demand and still be wrong. "Timer" and "interval timer" look similar and attract different people. If the searcher's intent does not match what your app opens to, a high rank converts badly and Apple learns that your listing is a poor answer for that query.

**2. Can you realistically place?** Look at who holds the top five. If they are all household names with hundreds of thousands of ratings, the term is a flag to plant later, not a term to build your title around now. If two of the five are apps roughly your size, that is a real opening.

**3. Is the demand plausible?** This is where difficulty and traffic scores belong — as a sanity check at the end, not as the starting filter. A score that says a term is easy while the top five are all incumbents is telling you the formula does not have enough evidence, not that the term is easy.

Track before you change, not after

The single most common mistake is changing metadata and *then* starting to measure. You now have no baseline, so whatever happens next is uninterpretable.

Track your candidate terms for at least a week before you touch anything. Ranks move on their own — Apple reindexes, competitors update, seasonality shifts. You need to know what your normal noise looks like before you can recognise a real change.

Then, when you do change metadata, record the date of the change against the rank timeline. Not because the timeline proves causation — it does not, and anyone who tells you otherwise is selling something — but because without the marker you cannot even form a hypothesis. A rank that climbed steadily for six days before your change was not caused by your change.

Keyword research is per-storefront, always

A term researched in the US tells you almost nothing about Germany. Different competitors rank, different words are used for the same concept, and Apple's autocomplete returns different completions. "Fitness tracker" in the UK is not the same market as "fitness tracker" in India, even though the string is identical.

If you operate in several countries, research each one separately or be explicit that you are guessing. Translating your US keyword set is the single most common localization error, and it produces keyword sets full of terms no local user would type.

What a good keyword record looks like

For each term you decide to track, you should be able to answer:

  • Which store, country and language was this observed in?
  • When was it last observed, and what was the rank?
  • How deep did the search go — top 50, top 100, top 200? A "not ranked" from a top-50 search is a different statement from a "not ranked" from a top-200 search.
  • Which apps were above it, and what do their titles target?
  • If a score is attached, what formula version produced it and what observations fed it?

If your tooling cannot answer those questions, the numbers it shows you cannot be audited, and an unauditable number is not a fact.

The workflow, condensed

  1. Harvest candidates from your own listing, autocomplete and the titles of apps outranking you.
  2. Filter by intent first. Drop anything that does not match what your app actually does.
  3. Check the top five for each survivor. Keep the ones where apps your size are placing.
  4. Track everything that survives for a week before changing a thing.
  5. Change metadata, record the date against the timeline, and keep tracking.
  6. Review after a full reindex cycle — not after two days.

None of this requires a proprietary dataset. It requires repeated public observations, stored with enough context to be checked later, and the discipline to keep the store's facts and your tool's opinions in separate columns.