Executive summary
Amazon S3 Vectors now “evaluates metadata filters before running similarity search, returning up to 5x more of the matching vectors when your filter is selective.” No extra cost, all commercial Regions plus China.
Read that backwards and it is a confession about the old behaviour. If pre-filtering returns five times
more matches, the previous mode was surfacing as little as a fifth of them — and
the documentation now says what that looked like:
“On a CLASSIC index, queries with filters may return fewer
than top K results when the vector index contains very few matching results.”
Fewer results, no error. That is precisely the failure #61 was about: a wrong answer from missing context is indistinguishable from a weak corpus.
What changed
There are now two index modes, and the difference is the order of operations.
ENHANCED — pre-filtering.
“S3 Vectors first identifies the vectors that match your filter, then only
searches those vectors for the most similar ones, ensuring high recall even when filters match a small
fraction of vectors.”
CLASSIC — evaluated in tandem.
“S3 Vectors searches through candidate vectors in the index to find the top K similar vectors
while simultaneously validating if each candidate vector matches your metadata filter
conditions.” The search decides the candidate set, and the filter then thins it.
And the cutover is a date, not a flag.
“Vector buckets created on or after September 30, 2026 create
ENHANCED indexes. In a vector bucket created before that
date, indexes have the index mode CLASSIC.”
So this is not a feature you wait for. Every S3 Vectors bucket that existed yesterday is on the old behaviour until you move it.
Architecture
The practical difference shows up when the filter is selective, which is the normal case for anything tenant-scoped, date-scoped or permission-scoped.
AWS's own example is a corpus of movie embeddings filtered by
genre='mystery': on ENHANCED, S3 Vectors
“first identifies the vectors where the genre metadata is 'mystery', then only searches those
for the closest matches.” On CLASSIC, it finds the nearest
neighbours across everything and keeps whichever happen to be mysteries. If mysteries are 2% of the
corpus, most of the candidate set is discarded and the result is short.
Ask for 10 chunks, get 3, and pass those 3 to the model. The model answers from three chunks, fluently, and nothing in the response says the retrieval was short. #61 argued the measurement to build first is "how often was the answer present in what came back" — on a CLASSIC index with a selective filter, that number could be low for a reason that has nothing to do with chunking, embeddings or the model, and nothing surfaces it. Document-level permissions are exactly this shape: a filter that matches a small share of the corpus, on every single query.
Twelve operators, and a new one worth having
The filter language covers $eq, $ne,
$gt, $gte, $lt,
$lte, $in, $nin,
$exists, $and, $or
— and now $startsWith, a
“Prefix match. Matches values that begin with the specified string.”
That is the one that makes hierarchical scoping practical:
{"s3_path": {"$startsWith": "/marketing/"}}. Previously a folder-scoped
search meant enumerating paths with $in.
It is also available without migrating:
“$startsWith requires ENHANCED query
behavior. It is available on a vector index whose index mode is ENHANCED, and
on a CLASSIC index when the request sets
queryMode to ENHANCED.” A
per-request override is the cheapest way to compare the two behaviours on real queries before changing an
index.
Business value
The value is recall on filtered queries, and filtered queries are most production RAG queries. Anything multi-tenant filters by tenant. Anything permission-aware filters by access. Anything time-bounded filters by date. All three are selective filters, which is exactly where the old mode lost the most.
It also costs nothing: “available at no additional cost in all commercial AWS Regions where Amazon S3 Vectors is available, and in the AWS China Regions”, with “no change to how you write vectors.”
Read alongside #63, it strengthens the case for S3 Vectors specifically. That post found the filterable metadata budget tight — 2 KB of a 40 KB total — and tight filterable metadata matters much more when filtering works properly.
Security considerations
This is a correctness issue for permission-filtered retrieval, and worth stating plainly.
If you implemented document-level access control as a metadata filter on a
CLASSIC index, the filter was never returning the wrong documents — it
was returning too few of the right ones. That fails safe rather than open, which is the right
direction, but it means results were silently incomplete for authorised users.
Check the inverse assumption, though. Anyone who compensated for short result sets by
raising top-K, widening filters or falling back to an unfiltered query will now get materially different
behaviour on an ENHANCED index. A fallback that quietly drops the filter is
the dangerous pattern here, and it is the kind of workaround short results invite.
And one irreversible decision to get right at index creation. “Once a metadata key is designated as non-filterable during index creation, it can't be changed to filterable later.” With filtering now substantially better, the filterable/non-filterable split is a more consequential choice than it was yesterday, and it is made once.
Cost considerations
No charge for the feature. The cost shows up as latency, and the documentation is unusually direct about the shape of it:
“Because an ENHANCED index evaluates your filter before it searches,
the work a filtered query does depends on the filter you send. Three things increase that work,
and with it the query's latency: a larger vector index, a filter that matches a larger share of
the vectors in the index, and a filter with more constraints.”
Note the middle one — a filter matching a larger share is more expensive. So pre-filtering
is fastest exactly where it helps most, and slowest where it helps least. A filter matching 90% of the
corpus does the most work for the least benefit, which is an argument for not filtering at all in that
case rather than for staying on CLASSIC.
Operational considerations
Find out which mode your indexes are in, by bucket creation date. Before 30 September
2026 means CLASSIC. That is a one-line audit and it tells you whether every
filtered query you have run was subject to the old behaviour.
Compare before migrating, per query, using queryMode. The
per-request override exists precisely so you can run your real queries both ways and measure the
difference rather than assume it. That is the honest version of adopting this.
The 100-constraint limit is new and applies only to ENHANCED.
“A single query filter can use up to 100 filter constraints. Each value the filter evaluates
counts as one constraint… This limit only applies to ENHANCED
indexes.” So migrating can break a query that worked — specifically a long
$in list, where each element counts. Anyone enumerating tenant IDs or document
paths in a filter should count them before switching.
Deployment is still in progress. AWS says it is “in the process of deploying this change and plan[s] to complete the deployment in the coming days”, so behaviour may vary by Region for a short while. Worth knowing before you conclude that a comparison test disagrees with the documentation.
Tradeoffs
Recall against predictable latency. CLASSIC does roughly the
same work regardless of the filter. ENHANCED does work proportional to what
the filter matches, so latency becomes filter-dependent.
Correct results against a new ceiling. The 100-constraint limit buys the pre-filtering, and it did not exist before.
A better filter language against an irreversible schema choice.
$startsWith and working pre-filters make the filterable set more valuable, and
non-filterable keys can never be promoted.
Implementation guidance
Audit bucket creation dates first. Anything created before 30 September 2026 is
CLASSIC, and its filtered queries have been returning short result sets.
Count your $in lists before migrating. Each element is a
constraint against the 100 limit.
Use queryMode: ENHANCED to A/B your real queries. Measure
recall on the questions you actually serve, which is the measurement #61 recommended and almost nobody
builds.
Review any fallback that widens or drops a filter. Those were written to compensate for a behaviour that is changing.
Replace path enumeration with $startsWith. Fewer constraints,
clearer intent, and it works on CLASSIC with the query-mode override.
Best practices
Treat a short result set as a bug, not as a thin corpus. That instinct would have caught this a long time ago.
Decide the filterable/non-filterable split deliberately at index creation. It is the one part of this that cannot be revisited.
Do not filter on something that matches most of the corpus. On
ENHANCED that is the most expensive filter for the least gain.
Who should act on this
Anyone running S3 Vectors with metadata filters, which is to say anyone running multi-tenant, permission-aware or date-bounded retrieval on it.
Act now if your filters are selective — a tenant id, an access list, a recent-documents window
— because that is the case where the old mode lost the most and the new one costs the least. Take
more care if your filters use long $in lists, since the 100-constraint limit is
new, or if your application has workarounds built around short result sets.
Key takeaways
- S3 Vectors now evaluates metadata filters before the similarity search, returning up to 5x more matching vectors on selective filters.
- Which means the old mode could have been surfacing as little as a fifth of them.
- The documentation now states it: on a
CLASSICindex, "queries with filters may return fewer than top K results" — no error. - Buckets created before 30 September 2026 are
CLASSIC. This is a cutover date, not a rollout you wait for. - New
$startsWithprefix operator, bringing the filter language to twelve operators. $startsWithalso works on aCLASSICindex by settingqueryModetoENHANCEDper request.- New limit: 100 filter constraints per query,
ENHANCEDonly — and each element of an$inlist counts as one. - Latency now depends on the filter: bigger index, broader match and more constraints all cost more.
- A filter matching a large share of the corpus is the most expensive and least useful case.
- Non-filterable metadata keys can never become filterable — decided once, at index creation.
- No additional cost, all commercial Regions plus China, and deployment was still completing at announcement.
Comments