A wireframe globe graticule with a tilted square map cell laid over it, four copper points inside the cell and one brick-red point breaking its lower edge, a short measuring line running from the centre point to the red one and a longer line to a point inside

DynamoDB vector search without embeddings

Update, 22nd September 2026. I compared the vector search against a single geohash cell, which isn't how anyone builds it. Reading nine cells finds the shop it missed, so the conclusion below is narrower than it looks - here's what changed.

Serverless Advocate #85 asks Lee Harding which AWS service he's most excited about, and he picks DynamoDB vector indexes. Not for embeddings, though. His point is that most people hear "vector" and think semantic search, when a vector index is "a general-purpose tool for any domain where you need to find 'nearest neighbours' in some n-dimensional space". Geometry, topology, geography, sensor fusion.

So I had a go at the geography case.

The setup is a store locator, which is about the most ordinary geography question going. Seven real places stand in for shops, five in and around Leeds and two over in York, each carrying its actual latitude and longitude. You stand somewhere and ask which ones are nearest.

A shop is about as plain as a DynamoDB item gets. Partition key, a name, two numbers:

{
  "PK":        { "S": "PLACE#p-101" },
  "name":      { "S": "Briggate" },
  "lat":       { "N": "53.79756" },
  "lon":       { "N": "-1.54169" },
  "openUntil": { "S": "18:00" }
}

No sort key, and the coordinates are ordinary numbers sat on the item. Nothing about that is searchable by distance yet, so the question is what you have to add to make it so.

Degrees aren't a space

You can't hand a vector index a latitude and longitude as two numbers, because degrees aren't a space. Cosine measures the angle your pair makes with the origin, and the origin here is just (0, 0). Nothing is measured from there, so two places on opposite sides of the world can score as basically identical. Euclidean is less daft and still wrong, because a degree of longitude is 111km at the equator and 66km up here in Leeds, so a circle in degree space is an ellipse on the ground. Neither is fixable by picking a different distance function. It's the space that's broken, not the metric.

What works is projecting onto the unit sphere, which is three lines of trig:

x = cos(lat) * cos(lon)
y = cos(lat) * sin(lon)
z = sin(lat)

That triple goes on the item as position, an ordinary list of numbers sitting beside the name and the opening hours, because there is no vector type in DynamoDB. A vector index over that attribute is what makes it searchable.

Now every shop is a point on a ball of radius one, and the straight-line gap between two of those points is 2 * sin(theta / 2), where theta is the angle between them at the centre of the ball. That climbs steadily from zero to pi and never doubles back, so ordering by it gives you the true great-circle ordering. Not an approximation of it. The score converts back to kilometres with 2 * 6371 * asin(score / 2), which across those seven shops agrees with haversine to about half a metre.

Next to a geohash

Worth running that next to the way proximity normally gets done in DynamoDB, which is a geohash. It interleaves the two coordinates into a single string, so nearby places end up sharing a leading prefix and "what's near me" becomes a prefix read. That's a lovely trick, and DynamoDB is very good at prefix reads.

Briggate's full hash is gcwfhct3. To make that queryable you take the first four characters as a cell, key a GSI on it, and keep the whole hash as the sort key so neighbours inside a cell order by position rather than by id:

{
  "PK":      { "S": "PLACE#p-101" },
  "name":    { "S": "Briggate" },
  "geohash": { "S": "gcwfhct3" },
  "GSI1PK":  { "S": "GEO#gcwf" },
  "GSI1SK":  { "S": "gcwfhct3#p-101" }
}

So "what's near me" becomes one Query on GSI1PK = GEO#gcwf, which is cheap and exact. It's also where the problem lives: the cell is the partition key. A shop in a neighbouring cell isn't ranked lower, it's in a different partition, and the query never goes near it.

I put the query point on the concourse at Leeds station, which sits in geohash cell gcwf, and asked both. Here are the five nearest shops of the seven, the York pair being a good 35km out and no trouble for the search or the cell query:

Shop Distance Cell In the gcwf prefix read?
Briggate 0.45km gcwf yes
The Headrow 0.48km gcwf yes
Kirkgate Market 0.53km gcwf yes
Holbeck 0.92km gcwc no
Hyde Park 2.24km gcwf yes

Holbeck is half a mile or so south of Leeds centre, and the cell boundary runs between it and the station. So the geohash hands you four shops, misses the one at 0.92km, and includes Hyde Park at two and a half times the distance because it happens to share four characters with where you're stood. That cell is roughly 23km by 19km, so this isn't a pathological case I've constructed: any query near an edge has it, which is most of them. Everyone who's built one of these knows about the boundary problem in the abstract. It's a bit sharper when the thing hands you a shop over twice as far away and looks perfectly happy about it.

The access patterns are loaded in the console below, if you'd rather press it than take my word for it:

That's the engine running in your browser on the seven shops above. Run the vector search, then switch to the cell query, and Holbeck is the row that stops coming back. There's a third run in there too, which is what the update at the end is about.

Three lines of trig is still a model

So it's tempting to call this the version with no model in it. No weights to ship, and nothing to re-embed when a library version moves.

That's not quite right. Something decided degrees weren't the space and the unit sphere was, and that decision is now part of the index's contract in exactly the way an embedding model would be. Store vectors built one way, query with vectors built another, and nothing will tell you. Swap two components of the projection, or move to a different reference sphere, and you still hand over three finite numbers. The dimension count still matches, and the dimension count is the only thing a vector index ever checks. Every result comes back looking perfectly reasonable.

Which is the same failure people hit with a mismatched embedding model, with no machine learning anywhere near it. So write down whatever turns your data into numbers, keep it next to the index, and version it. When it changes, re-derive the lot.

Have a go

The full write-up for the store locator is on accesspatterns.dev, and there's a geohash-only version of the same seven shops if you want to poke at the prefix read on its own. Everything runs on dynoxide's WASM engine, so there's no account and nothing goes near AWS.

I did two of Harding's other cases the same way if you want them. Sensor fusion is five sensor channels on wildly different scales, where the numbers need normalising before the distance means anything at all, and it's the clearest illustration of the point above. Design tokens matches a pasted hex to the nearest approved colour, where converting to CIELAB makes the distance function a consequence rather than a choice.

And if you've got a geohash doing proximity in production, it's worth firing a few queries near a cell edge and looking at what comes back. If you're reading one cell, it won't tell you it's missing anything. If you're reading nine, the thing to check is the one in the update below: whether your cell size actually guarantees the radius you advertise.

Update: the comparison wasn't fair

Roger Chi pointed out on LinkedIn that I'd compared the vector search against a single geohash cell, and nobody builds it that way. He's right, and it's the weakest thing in the piece above. Reading one cell is the quick version, and picking it made the vector index look better than it is.

The normal way is to read the query's cell and its eight neighbours, put the results together, then work out the real distances and drop anything over five kilometres. Do that and it finds Holbeck. You get the same five shops in the same order as the vector search.

I also said the vector index can't do a radius, and that was too strong. The score converts back into kilometres, so you can filter on it. It won't give you everything, though. The search returns a hundred results at most and there's no second page, so if four hundred shops sit within five kilometres you'll get a hundred of them and nothing tells you the rest exist.

What nine cells cost

Nine requests instead of one, and 4.5 read units instead of 0.5, going by what the engine on this page meters. They go out in parallel, so it's still about one round trip.

Six of the nine come back empty and you pay for them anyway. An empty Query is charged as if it read 4KB, halved because reads on a GSI are eventually consistent.

Picking the cell size

Nine cells cover less ground than you'd think. Your query point can sit anywhere in the middle cell, including right at its edge, so the distance you're guaranteed to reach in every direction is only the short side of one cell. That's well known. I just hadn't put numbers on it before.

characters cell (N-S × E-W) guaranteed reach empty calls, this seed
4 19.5 × 23.1km 19.5km 6 of 9
5 4.9 × 2.9km 2.9km 7 of 9
6 0.6 × 0.7km 0.6km 8 of 9

Four characters covers the five kilometres I'm filtering to. Five characters doesn't. On these seven shops it happens to give the same answer for the same price, but it only guarantees 2.9km, so a query point a few hundred metres away could miss a shop at 4km and you'd never know. Covering 5km at five characters would take twenty-five cells, not nine.

Six characters only guarantees 600m. It drops Holbeck at 0.92km and Hyde Park at 2.24km, so nine requests get you three shops where one request got you four.

I've used nine cells throughout because it's the easiest version to follow. It isn't the best one: dynamodb-geo and H3 work out how many cells your radius actually needs, so nine is a fact about the simple approach rather than about geohashes.

So the claim I published was too broad. A vector index doesn't beat a geohash at finding the nearest thing, and the boundary problem I spent the article on is fixable - nine cells fix it, for nine requests.

The cell size is the part you're stuck with. It lives in the partition key, so you pick it before you write anything, and the same choice has to work for Briggate and for the moors above it. That's the real difference between the two. A geohash gives you whatever happens to be in the area, which is everything in a city centre and nothing on the moors. The vector search always gives you your five, wherever you are.