Databricks Lakebase Search Challenges the Separate Search Stack
Databricks has made Databricks Lakebase Search generally available on AWS and Azure, putting two search engines inside its managed Postgres service. The September 28 release supports vector retrieval and BM25 keyword ranking without requiring a separate search database. That challenges a familiar AI architecture: Postgres holds operational records, while another system indexes copies for retrieval.
The company says its new vector extension can search 100 million vectors with 97% recall and 71-millisecond P99 latency. It also claims twice the throughput of the next-best system in its benchmark and four times lower cost than cloud Postgres running pgvector. Those results are substantial, but Databricks produced the tests and did not publish a complete independent validation.
The larger story is not another vector index. Databricks Lakebase Search is an attempt to make operational Postgres handle semantic, keyword, and hybrid retrieval at agent scale. If the architecture works under real production workloads, some teams can remove a search service and the data pipelines surrounding it. The pressure falls on both pgvector deployments and dedicated search systems that justified their complexity through superior scale.
Databricks Lakebase Search Moves Retrieval Into Postgres
The release turns search from an attached service into a managed capability of the operational database.
The technical announcement introduces two Postgres extensions. lakebase_vector handles approximate nearest-neighbor search, which finds vectors close to a query without comparing every possible record. lakebase_text provides BM25, a ranking method that weighs term frequency, document length, and term rarity across a collection.
Both extensions are generally available for Lakebase projects on AWS and Azure. Developers can install either extension or combine them for hybrid retrieval. The combination matters because vector and keyword search solve different failure modes.
Vector search compares embeddings, which are numerical representations of meaning. It can match a query such as “fast sports car” with records mentioning an automobile model, even when those exact words are absent. Keyword search remains better for identifiers, names, error codes, product numbers, and other terms whose literal form carries meaning.
Hybrid search runs both methods and merges their rankings. An AI support agent might use semantic similarity to find conceptually related incidents while preserving an exact match for a specific error code. An ecommerce agent could interpret a shopper’s intent without losing a requested model number.
These operations run beside transactional records instead of against a separately synchronized copy. A developer can filter retrieval using current fields such as tenant, inventory status, access rights, or workflow state. Databricks says lakebase_vector applies filters while scanning index blocks, reducing the need to retrieve a broad candidate set and discard unauthorized or irrelevant rows afterward.
That design targets a persistent problem in retrieval systems. The freshest version of a record often lives in the application database, while the searchable version arrives later through an extraction pipeline. Even a short delay can expose an agent to deleted documents, outdated permissions, or inventory that no longer exists.
Keeping retrieval close to operational data reduces that synchronization window. It can also reduce the number of systems that engineers must monitor, secure, and repair. The change is especially relevant for teams building a searchable knowledge base, where access controls and document changes must remain aligned with search results.
Lakebase Search does not eliminate every data movement step. Embeddings still have to be generated, source content may originate outside Postgres, and lakehouse tables require synchronization before serving. The difference is that applications can query the resulting indexes through familiar Postgres types and operators.
Databricks is also tying the feature to its broader lakehouse platform. Its product documentation explains how Unity Catalog tables can be synchronized into Lakebase. During that process, an embedding column can become a Postgres vector, while source text can become a tsvector, PostgreSQL’s optimized representation for text retrieval.
The immediate change is therefore concrete. Lakebase now offers native managed indexes for meaning-based and exact-term search, and applications can query them alongside operational fields. The tension begins with what this consolidation replaces.
AI Agents Put the Separate Search Pipeline Under Pressure
Agent workloads make synchronization errors and idle infrastructure harder to justify.
A traditional search architecture usually contains at least two data stores. Postgres records transactions and application state. A search engine or vector database receives transformed copies through an extraction, transformation, and loading pipeline.
That split can work well at scale, but it creates operational obligations. Teams must detect failed updates, replay missing records, coordinate schema changes, preserve deletion semantics, and reproduce database permissions inside another system. They also need a plan for rebuilding indexes without interrupting the application.
AI agents amplify those obligations because retrieval becomes part of a decision loop. A conventional search page can tolerate an imperfect result while a user reviews alternatives. An agent may take action immediately after retrieving a record, making freshness and authorization more consequential.
An account-management agent illustrates the problem. It might search meeting notes semantically, match an exact contract identifier, and filter results by the current user’s permissions. If those three signals live in different systems, the application must reconcile them before the model can respond safely.
The same issue appears in commerce. A shopping agent can interpret an ambiguous request through embeddings, but availability and regional restrictions come from rapidly changing operational columns. Searching a stale copy can produce a convincing answer for an unavailable item.
Bursty usage creates another source of pressure. Human-facing enterprise search often follows predictable working hours. Agents can generate many parallel retrieval calls while planning, verifying, and revising a task. A single user request may trigger several searches rather than one.
Databricks designed Lakebase Search around this uneven demand. Lakebase separates durable storage from compute, keeping data in object storage while using memory and local NVMe as caches. Search compute can suspend when idle and resume when another query arrives.
The company reports a P90 first-query latency of 1.13 seconds after scaling to zero on an index containing 100 million vectors with 768 dimensions. It also says the same collection can be served with one Lakebase Compute Unit. These are company measurements, not universal expectations, but they show the intended operating model.
A one-second cold query will not suit every interactive application. However, it may be acceptable for an infrequently used internal agent if it avoids continuously running a large search cluster. Teams can keep compute active when latency matters and allow quieter environments to suspend.
Index construction also moves away from the primary transaction path. Databricks says it can train centroids from a sample, distribute vector assignment and quantization, then write independent index blocks. Future offloading to distributed engines such as Spark is part of the company’s direction, although the announcement says to stay tuned for that broader capability.
This matters because large index builds compete with transactional workloads when they consume the same processor, memory, and storage resources. Moving that work away from the primary database can reduce interference. It also changes the cost model from maintaining a permanently provisioned index server toward paying for active retrieval compute and durable storage.
The pressure target is not every dedicated search deployment. Large search teams often need specialized analyzers, custom ranking pipelines, advanced observability, or features developed over years. Lakebase instead pressures the common architecture where a second system exists mainly because Postgres search stopped scaling comfortably.
That distinction keeps the announcement grounded. Databricks is not arguing that one database should perform every search workload. It is arguing that more AI applications can postpone, simplify, or avoid the split.
Lakebase Vector Search Takes Aim at pgvector’s Memory Model
The primary contest is storage-backed Lakebase Search versus memory-heavy pgvector indexes at large scale.
Pgvector made Postgres a practical starting point for semantic retrieval. It adds vector types, distance operators, exact search, and approximate indexes without forcing developers into an unfamiliar database interface. It remains open source and widely available across hosted Postgres services.
Its standard approximate options include HNSW and IVFFlat. HNSW creates a multilayer graph connecting nearby vectors. It offers a favorable speed and recall balance, but graph construction takes time and the index consumes significant memory. IVFFlat groups vectors into lists and searches the most promising groups, reducing memory and build costs while generally offering lower query performance.
The project’s own pgvector guidance documents these tradeoffs. It notes that HNSW indexes build much faster when the graph fits within maintenance_work_mem. It also warns that increasing search candidates improves recall at the cost of query speed.
Databricks argues that these constraints become more difficult when an HNSW graph grows beyond one machine’s memory. Fetching a chain of graph nodes from remote object storage can produce many small, random reads. A design optimized for resident memory becomes less efficient when the working set is cold.
Lakebase vector search uses hierarchical inverted-file clustering to change that access pattern. Vectors are grouped into contiguous blocks. A query first scores cluster centroids, then reads blocks associated with the most promising clusters.
The extension combines that layout with RaBitQ binary quantization, which compresses each vector to roughly one bit per dimension for initial candidate scoring. Databricks describes this representation as about 32 times smaller than a standard 32-bit floating-point vector. The system then reranks a limited candidate set using full-precision vectors.
This mechanism makes the index friendlier to both object storage and local caching. A cold query reads several relevant blocks rather than following hundreds of graph links. A warm query can scan compact binary codes while keeping a smaller active footprint in memory.
Databricks says a single lakebase_ann index can hold more than one billion vectors. Its documentation also claims index builds are 50 to 100 times faster than HNSW. These figures describe the company’s implementation and should not be applied automatically to every schema, embedding model, or filter distribution.
The headline benchmark used 100 million vectors from the LAION dataset. According to Databricks, Lakebase delivered twice the throughput of the next-best tested system. It reported 97% recall at 71 milliseconds P99, meaning 99% of measured queries completed within that latency while retrieving the true neighbors at the stated rate.
The benchmark also produced the four-times-lower-cost claim against an unnamed cloud Postgres vendor using pgvector. Databricks notes that pgvector and DiskANN were tested on single large instances. That caveat limits the comparison because architecture, configuration, hardware, concurrency, and pricing assumptions can materially change results.
A benchmark can show that an approach deserves evaluation without settling the purchasing decision. Databricks has not established that every pgvector workload should migrate. Smaller indexes may fit comfortably in memory, and an existing pgvector installation can be inexpensive, portable, and easy to operate.
Pgvector also supports binary quantization, half-precision indexing, iterative scans, partitioning, and configurable search effort. Teams with tuned deployments have more options than a simple baseline chart suggests. The open-source extension works across many Postgres environments, while Lakebase Search belongs to a managed Databricks service.
Compatibility narrows the migration cost, however. Databricks says lakebase_vector uses pgvector’s vector types, distance operators, and query syntax. An application can retain familiar SQL while creating a lakebase_ann index instead of an HNSW or IVFFlat index.
That is a deliberate competitive move. Databricks is not asking developers to abandon the pgvector programming model. It is offering a different storage and indexing engine beneath much of the same interface.
Customer evidence provides one practical signal. Conexiom told Databricks that it runs hybrid BM25 search across more than 100 million rows with half the compute footprint of its previous pgvector setup. The account is useful because it describes an operational workload, but it remains a vendor-selected customer statement without independently published methodology.
The case for Lakebase vector search is strongest when the collection is large, query demand is irregular, and operational filters matter. The case becomes weaker when teams prioritize infrastructure portability, have predictable always-on demand, or already meet latency targets with pgvector.
Native BM25 Changes the Full-Text Search Equation
The quieter part of the release may be more important than the vector benchmark.
Many AI search products overemphasize embeddings. Semantic matching helps when users and documents express the same idea with different words. It is less dependable when a query contains an exact identifier that an embedding model treats as weak or unfamiliar.
Consider an agent searching for “CVE-2026-1234,” a customer account number, or a specific component name. Similarity search may return conceptually related records while missing the exact string’s importance. Keyword ranking supplies a separate retrieval signal that preserves literal matches.
Lakebase’s lakebase_text extension adds a lakebase_bm25 index compatible with PostgreSQL tsvector values and text-query operators. BM25 incorporates collection-wide term frequency and document length, helping rare terms contribute more than common ones.
PostgreSQL already provides substantial full-text functionality. It can parse documents, normalize words, remove stop words, build GIN indexes, and rank results with ts_rank or ts_rank_cd. The official ranking documentation notes that its built-in rank functions use lexical frequency, proximity, and structural information.
Those functions do not use global collection statistics in the same way as BM25. That difference matters when a product needs search-engine-style relevance rather than simple matching. Teams have historically added custom ranking logic or moved text into a dedicated engine.
Databricks says lakebase_text uses Block-Max WAND for top-K retrieval. This algorithm skips regions that cannot produce a result competitive with the current highest scores. Instead of fully scoring every matching document, the engine concentrates work on candidates that can enter the requested result set.
The approach complements vector retrieval. A support query could run against lakebase_bm25 for exact error text and lakebase_ann for semantically similar incident descriptions. Reciprocal rank fusion can then combine both ordered lists without assuming that their raw scores share the same scale.
This is where Databricks Lakebase Search becomes more than a faster vector index. It offers a search stack with two distinct retrieval models inside the same database. The operational row, the embedding, the text representation, and the filtering attributes can remain together.
That consolidation affects security as much as convenience. An application can express tenant boundaries and permission checks as SQL predicates alongside retrieval. Engineers still need to test whether every index path enforces filters correctly, but they avoid recreating an entire authorization model in a separate service.
It also simplifies write behavior. A newly inserted record can become searchable without waiting for a second database to acknowledge an event. Updates and deletes remain within a familiar transactional environment, although index-maintenance timing and synchronized lakehouse sources still require measurement.
Dedicated engines retain important advantages. Elasticsearch and similar systems support broad language analysis, customized scoring, aggregations, highlighting, query tooling, and operational controls developed specifically for search. Lakebase’s BM25 support does not erase those differences.
The meaningful comparison is therefore architectural. If an application needs semantic retrieval, exact-term ranking, fresh operational filters, and ordinary SQL, Lakebase can cover more of that workload in one place. If search itself is the product, specialized capabilities may still justify a separate system.
The Benchmark Leaves Production Questions Unanswered
Databricks has shown an attractive mechanism, but buyers still need workload-specific evidence.
The largest uncertainty is benchmark independence. Databricks selected the systems, configurations, dataset, instance shapes, and cost assumptions behind its published comparison. The company identifies VectorDBBench and the LAION 100 million dataset, but the announcement does not provide enough detail to reproduce every result from the article alone.
Recall and latency also interact. Approximate retrieval intentionally avoids exhaustive comparison, so engineers tune how many clusters or candidates a query examines. Higher recall often demands more work. A single performance point cannot describe the full curve across different target recall levels.
Filtering can change that curve again. Real business queries may restrict results by tenant, geography, time, inventory state, or authorization. A uniformly distributed benchmark does not necessarily represent highly selective or uneven production filters.
Data shape matters too. Image embeddings from LAION differ from enterprise document embeddings, product catalogs, source code, or customer records. Dimensions vary, duplicates appear, updates arrive unevenly, and some tenants dominate traffic. Each factor can affect cache behavior and index quality.
Cold-start performance deserves careful interpretation. The reported 1.13-second P90 applies to a specific 100-million-vector, 768-dimension configuration. Applications with strict interactive targets may need active compute rather than scale-to-zero. Teams should test both the first query and the following burst.
Operational constraints also need attention. Enabling Lakebase Search restarts every compute resource in a project, drops active connections, and cannot be reversed, according to the documentation. That makes activation a planned infrastructure change rather than a harmless extension toggle.
Portability is another tradeoff. Lakebase presents standard Postgres types and familiar pgvector syntax, but its new index access methods are proprietary managed capabilities. A team can retain much of its application SQL while still becoming dependent on Databricks for index behavior, scaling, and pricing.
The same concern applies to BM25. Standard tsvector columns remain recognizable Postgres objects, but the lakebase_bm25 index and its execution characteristics are specific to Lakebase. Moving away may require rebuilding indexes and retesting ranking quality elsewhere.
Cost claims require direct measurement. Serverless suspension can lower expense for irregular use, yet high sustained concurrency may favor a different model. Embedding generation, synchronized tables, storage, data transfer, and surrounding Databricks services contribute to the total architecture.
Teams should therefore evaluate Lakebase Search with representative questions rather than a generic leaderboard. A useful test corpus includes current records, deleted records, access-controlled documents, rare identifiers, ambiguous natural-language queries, and the filters most likely to reduce recall.
They should also compare operational outcomes. Measure data freshness, failure recovery, index-build impact, permission consistency, and the staff time required to manage pipelines. Removing an external service can be valuable even when raw query latency changes little.
None of these questions invalidates the release. They define what “state of the art” must mean outside a vendor benchmark. The architecture has a credible technical rationale, but production evidence must show that its advantages survive each buyer’s data distribution and workload.
What to Watch After Lakebase Search Reaches GA
Three signals will determine whether Lakebase Search becomes a default Postgres feature or remains a Databricks-specific option.
The first signal is reproducible performance. Independent tests should compare Lakebase with tuned pgvector, DiskANN-based services, and dedicated search engines across several recall targets. They should publish instance specifications, concurrency, filter selectivity, cache state, index-build time, and complete cost assumptions.
Results close to Databricks’ claims would strengthen the case that storage-backed clustered indexes fit large serverless collections better than memory-oriented graphs. A wide gap would weaken the performance narrative, even if consolidation still offers operational benefits.
The second signal is adoption among teams replacing two-system architectures. Conexiom supplies an early example, but the market needs more accounts describing production scale, update frequency, query volume, and permission models. The most persuasive stories will document a removed search cluster or ETL pipeline, not simply a successful demonstration.
Adoption will also reveal whether familiar Postgres syntax reduces migration friction. If teams can change index definitions while keeping their data model and queries, Lakebase Search has a practical route into existing applications. If migrations require extensive ranking changes, the compatibility claim will carry less weight.
The third signal is competitive response. Pgvector continues to add options for quantization, filtering, and iterative scans. Managed Postgres vendors can improve storage architecture or introduce their own search extensions. Dedicated search providers can emphasize mature ranking controls, deployment flexibility, and hybrid retrieval features.
Databricks began positioning Lakebase as managed Postgres for AI applications in its 2025 launch announcement. This release makes that positioning more concrete. Transactions alone do not make a database agent-ready if every serious retrieval query still leaves the system.
The immediate takeaway is narrower and more useful. Databricks Lakebase Search gives developers one managed place for operational records, vector similarity, BM25 ranking, and SQL filtering. Its clustered, quantized vector index directly addresses the memory model that constrains large pgvector deployments.
The next decision belongs to engineering teams. Build a representative retrieval test, include cold and warm traffic, apply real permission filters, and compare the complete operating burden. If Lakebase preserves relevance while removing synchronization infrastructure, the architectural simplification will matter more than any single benchmark bar.



