Skip to content
Try Raven →
All posts
ExplainerBy the Raven team10 min read

How AI Guesses Where a Photo Was Taken

Landmarks, license plates, the angle of the sun. A plain-English tour of the visual clues an AI reads to work out where on earth a photo was shot.

Short answer

AI guesses where a photo was taken by reading visual evidence: architecture, the script on signs, vegetation, road markings, vehicles, terrain and the angle of the light. A multimodal model such as Google Gemini weighs those clues together and returns a probable region. Raven's results are entertainment-only estimates and can be wrong.

Abstract topographic map with faint contour lines and thin teal analysis vectors.

You have seen the photo. A friend's holiday snap turns up in your feed: a sun-drenched street corner, a small cafe, unfamiliar trees. No caption, no tag. Your mind starts hunting for clues almost before you decide to. Is that architecture Spanish? Are those characters on the sign Cyrillic? What sort of car is that? You are playing geographic detective, and it is a game that image models have become surprisingly good at.

Raven does the same thing on demand. You upload a picture, Google's Gemini model reads the pixels, and a probable location comes back with a short explanation. It can feel like a trick, but there is no trick in it. It is observation and inference — the digital version of noticing the particular clay on a visitor's boots. Every photo is a case file, and most of them contain far more evidence than the person who took the picture realised.

What does an AI actually look at in a photo?

It looks at everything visible at once: building materials and rooflines, the script and language on signs, plant species, soil and rock colour, road markings and sign shapes, vehicle types and number-plate formats, and the hardness and colour of the light. No single clue decides the answer; the combination does.

The obvious clues are the big ones. A shot of the Eiffel Tower is Paris, and the model will say so in a fraction of a second. But famous buildings are a rounding error in the world's photographs. Fewer than 1,300 places carry a UNESCO listing, and the vast majority of pictures ever taken contain none of them. The real skill lies in the mundane patterns that quietly define a place: the steep gabled roofs and half-timbering of the Rhine valley, the flat white parapets of the Aegean, the tin-roofed verandas of the tropics.

Language is the other heavyweight. Even when the model cannot read a word, it recognises the writing system. The looping ascenders of Thai, the geometric blocks of Korean Hangul, the connected flow of Arabic — each one removes most of the planet from consideration in a single step. It does not need a clean sign, either. A blurred menu behind glass or a scrap of graffiti can be enough to catch the shape of a letterform and add it to the pile. There is a longer catalogue of these signals in our field guide to the ten visual clues AI uses to find a location.

Beyond the built environment, the natural world leaves its own fingerprints. Vegetation is really a proxy for climate: an olive tree implies a Mediterranean summer drought, a fir implies a cool temperate belt, a mangrove implies a warm coast. The Koppen climate classification is the formal version of the same intuition, and it is worth noticing what that means in practice — plants narrow a photo to a climate band that wraps around the globe, not to a country. Soil colour works the same way. Deep red laterite suggests the tropics or the Australian interior; pale chalky ground suggests limestone country.

Light is the quietest evidence in the frame and often the most physical. The height of the sun at midday depends on latitude and the date, so shadow length is a genuine measurement rather than an impression. Hard, short shadows and a deep blue sky suggest a low latitude or high altitude. Long, soft shadows and a flat white sky suggest somewhere further from the equator, or simply a cloudy afternoon that has erased the clock entirely.

Which clues carry the most weight?

The strongest clues are the ones a country legislates rather than chooses by taste: which side of the road traffic uses, number-plate shape and colour, road-sign geometry, and the script on public signage. Those are enforced nationally, so they slice the map cleanly where architecture and plants only shade it.

Some of the most decisive evidence is the sort nobody notices while living inside it. The systems that govern movement are written into law, which makes them unusually clean signals. The single most powerful is which side of the road the traffic uses. Left-hand traffic covers roughly 75 countries and territories, including the United Kingdom, Ireland, Japan, India, Australia and much of southern and eastern Africa. Spot it and you have removed most of the Americas and nearly all of continental Europe in one move, as the survey of left- or right-hand driving makes clear.

  • Side of the road. A binary, legally enforced split. The cheapest and most reliable filter available.
  • Number plates. Rarely legible, but the proportions and colour block are diagnostic: the tall narrow European plate with its blue band, the wide North American plate, the yellow rear plate of the UK and the Netherlands.
  • Sign geometry. Countries that signed the 1968 Vienna Convention share sign shapes and colours; the United States, with its own manual, does not. Warning signs alone separate large parts of the world.
  • Script and language. A writing system is a hard boundary. Even a partial word narrows the field further, separating Portuguese from Spanish or Dutch from German.
  • Utility furniture. Bollards, kerb paint, manhole covers, bin design, overhead cabling and pole types are set by municipal standards and vary sharply between neighbouring countries.

Notice the pattern. Anything decided by a national standards body is a strong clue; anything decided by taste, fashion or the global supply chain is a weak one. A plastic garden chair tells you nothing, because the same chair is sold on six continents. A kerbstone painted in a specific two-colour scheme can be worth more than the entire building behind it.

How does the model turn weak clues into one guess?

It accumulates them. Each detail shifts the odds slightly rather than settling the answer, and the region where the most independent clues overlap becomes the guess. That is why a photo full of individually vague evidence can still yield a confident result, and why one strong clue can be outvoted by several weak ones.

Think of it as weighing evidence rather than looking things up. Terraced brick housing is consistent with Britain, but also with parts of the Netherlands and the American northeast. A grey overcast sky is consistent with half the temperate world. Traffic on the left removes the Netherlands and the United States. A yellow rear number plate keeps Britain and returns the Netherlands to contention. A red pillar box settles it. Individually every one of those observations is weak. Stacked, they converge on a single country.

This is possible because the model is multimodal — it holds the image and the question in one shared representation rather than running a vegetation detector and a signage detector separately and stapling the outputs together. If that idea is new, what multimodal AI means in plain English is the shortest route into it. The practical consequence is that the model can explain itself, and the explanation is usually more interesting than the pin on the map.

The same accumulation logic explains the failures. When every visible clue is of the shared, portable kind — a whitewashed wall, an olive tree, hard midday sun — the evidence genuinely does point at several countries at once, and a careful answer widens rather than narrows. We wrote about that case separately, in how AI handles photos with multiple possible locations.

What is the AI not doing?

It is not reading EXIF metadata, searching a database of your photos, matching against a private image index, or accessing any location service. Raven sends the image to the model, receives a text answer, and discards the picture. The guess is inference from pixels alone, with nothing else to draw on.

This matters because people reasonably assume something more invasive is happening. It is not. There is no lookup against a library of your holidays, no reverse image search against the open web, no GPS tag being quietly harvested. If you strip a photo of every scrap of metadata and upload it, the answer is unchanged, because the metadata was never used. The uploaded image is processed in memory, passed to the model, and gone when the response returns — never written to disk, never stored in a bucket, never kept in a database.

It is also not an authority. The model has no way to verify its own answer, no second source to check against, and no notion of whether it has seen this scene before. What it has is a very large amount of pattern-matching and a fluent way of describing it, which is a combination that reads as more certain than it is.

How accurate is it, and when does it fail?

Accuracy depends almost entirely on the photograph. Frames with legible signage, distinctive infrastructure or a known landmark can reach the right city. Interiors, tight crops, chain stores and flat overcast light often support only a continental guess. A confident-sounding answer is not the same as a correct one.

The honest summary is that the spread is enormous and it is driven by the picture, not the model. Hand it a street scene in Lisbon with tiled facades, a tram and a legible shopfront and the answer is likely to be specific and right. Hand it a close crop of a plate of food in a windowless room and no system on earth can do better than a shrug. The catalogue of hard cases is worth knowing before you draw conclusions about a wrong answer: what makes a photo hard to geolocate covers interiors, chain environments, heavy filters and tight crops.

The confidence figure attached to a result deserves the same scepticism. It reports how decisively the model's own scoring favoured one answer over the alternatives it considered — not the probability that the answer is true. A high number on a photograph with thin evidence is a warning sign rather than a reassurance, which is the argument in how much to trust an AI confidence score.

A worked example: one grey street corner

Imagine an unremarkable photograph: a road, some houses, a parked car, no people, no landmark, flat grey light. Start with the traffic, which is on the left — most of the Americas and continental Europe are out. The car has a yellow rear plate, narrowing to a short list that includes the UK and the Netherlands. The Netherlands drives on the right, so it is gone. The houses are two-storey brick terraces with bay windows and chimney stacks, which is a British vernacular. The kerb has a double yellow line rather than a solid painted edge. A wheelie bin sits on the pavement in a shade of green a council chose. The road sign at the corner is a white rectangle with a black border and a place name in a distinctive typeface.

None of that required a landmark, a caption or a single readable word beyond a place name. Seven ordinary details, each nearly worthless on its own, converged on a country and then a region. That is the whole mechanism, and it is exactly what the model is doing — faster, across more categories at once, and with a much larger memory of what ordinary streets look like in places it has never been.

Which is the genuinely interesting part of all this. Not that a machine can find a landmark, but that the world turns out to be legible in its dullest details. Every place has a visual dialect, written in kerbstones and roof pitches and the particular green of municipal paint. Once you start reading it, it is difficult to stop.

Upload a photo and see which clues Raven picks out first.

Upload a photo →

If you would rather see it than read about it, the AI photo location finder runs exactly the process described above on a photo of your own.

Frequently asked questions

Does the AI read the GPS coordinates stored in my photo?
No. Raven works from the picture itself, not from the file's metadata. The guess comes from what is visible in the frame, which is why a screenshot or a re-saved copy with no location data still gets an answer.
How accurate is an AI location guess?
It varies enormously with the photo. A frame containing a legible street sign or a famous building can land on the right city; a bare beach or a hotel corridor may only support a continent-sized guess. Treat every answer as an estimate.
Is my photo stored anywhere after the guess?
No. The image is held in memory only for as long as the request takes, passed to the model, and discarded when the response is returned. It is never written to a disk, a bucket or a database.
Can I use this to work out where someone else is?
No, and please do not try. Raven is built for curiosity and entertainment. It has no access to private data, its output is frequently wrong, and using it to track a person is a misuse of the tool.

Sources

  1. Left- and right-hand trafficWikipediaAround 75 countries and territories drive on the left, which is why the side of the road is such a strong first filter.
  2. Vienna Convention on Road Signs and SignalsWikipediaOpened for signature in 1968; it standardised sign shapes and colours across much of Europe, which makes non-signatory countries stand out.
  3. World Heritage ListUNESCO World Heritage CentreMore than 1,200 inscribed sites — the small set of buildings and landscapes an image model can name outright.
  4. Koppen climate classificationWikipediaThe standard scheme behind the idea that vegetation and light imply a climate band rather than a country.

Reminder

Raven is built for entertainment and curiosity. Its guesses are AI estimates that can be wrong, and it must never be used to track or identify real people. Uploaded photos are processed in memory and immediately discarded — never stored.

Get Geospy AI for iPhoneDownload free