top of page

Cloudflare Is Introducing Web Search API via AI Gateway, Turning Search Into Infrastructure

3 hours ago
14 min read

Cloudflare is introducing Web Search API via AI Gateway with three search providers, moving live web retrieval into the same control plane as model inference. The beta supports Ceramic.ai, Exa, and Linkup through one interface. Developers can call it from a backend, a Cloudflare Worker, or an agent workflow.

The important change is not that another web search endpoint exists. Cloudflare is placing search beside model routing, logs, security controls, credentials, and usage management. That positioning turns retrieval from an isolated integration into managed AI infrastructure.

The move also creates a clear contest. Developers can use search tools built into model platforms, connect directly to specialist search companies, or put retrieval behind an independent gateway. Cloudflare is betting that teams want the third option, especially when one application uses several models and search providers.

Introducing Web Search API via AI Gateway Changes the Control Point

Cloudflare now gives developers one managed route to three independent search services, without tying retrieval to a particular language model.

The company announced the beta on October 2, 2026. According to the launch announcement, requests can use Ceramic.ai, Exa, or Linkup. Ceramic.ai becomes the default when an application does not specify a provider.

Every response follows a shared structure containing a title, URL, and description for each result. Metadata can also include the query, a request identifier, and latency information. That consistency matters because provider-specific differences often leak into application code.

Developers can access the service through a standard REST endpoint. Cloudflare Workers users can instead call env.AI.websearch() through an AI binding. Both methods send the request through an existing AI Gateway.

The REST route accepts a query, provider name, result limit, and gateway configuration. The Workers binding exposes equivalent controls through JavaScript or TypeScript. Cloudflare’s implementation guide says a query can contain up to 1,024 characters, while one request returns up to 10 results.

That limit shows what the product is designed to do. It is a context retrieval layer for model calls, not a replacement for a conventional search results page. An application gathers a focused set of sources and places relevant snippets into a model’s working context.

The service supports two credential paths. Teams can spend their AI Gateway credits, or they can store a provider key and select it through a bring-your-own-key alias. Cloudflare retrieves that credential inside the gateway rather than requiring the application to transmit it with every search request.

This architecture gives teams a common boundary for authentication. It also reduces the number of external credentials distributed across applications, deployment systems, and developer machines. A compromised application token still presents risk, but credential sprawl becomes easier to contain.

Cloudflare says web search requests appear inside normal AI Gateway observability logs. Teams can inspect request activity alongside model traffic rather than operating a separate monitoring stack. Access policies can also determine which applications reach particular search providers.

That integration creates the article’s central tension. A direct search API offers fewer intermediaries, while a gateway offers more operational control. Cloudflare must prove that the control layer saves more complexity than it introduces.

Search Is Becoming Part of the AI Gateway Stack

The competitive pressure falls on model platforms and specialist search vendors because Cloudflare is separating live retrieval from the model that consumes it.

Many models already offer native web search. Cloudflare’s existing gateway documentation lists supported search tools from OpenAI, Anthropic, xAI, and Alibaba. Those tools remain attached to their respective provider interfaces and model capabilities.

Native search can be convenient when a team commits to one model family. The model decides when to search, the provider formats the evidence, and the same service produces the answer. That path can minimize orchestration work for a straightforward assistant.

It becomes less convenient when an application changes models. Tool schemas, supported models, citation formats, regional availability, and retention terms can differ. A team may need separate implementations for each provider, even when every path performs the same basic retrieval task.

Cloudflare Web Search API changes that boundary. Search becomes an application-controlled step with a common response format. The retrieved material can feed a model available through Workers AI, a model routed through AI Gateway, or another inference service.

This separation matters for agents, which often perform several searches before producing one answer. A research agent might begin with a broad query, identify a company or document, and issue narrower follow-ups. Those calls need predictable logging and permissions because they can outnumber the final model request.

The gateway approach also supports explicit provider selection. A developer can route one workload to Ceramic.ai and another to Exa or Linkup. The application does not need to change its overall integration whenever the selected provider changes.

Cloudflare describes each provider as serving a different retrieval profile. Its provider documentation says Ceramic.ai operates an independent index exceeding 40 billion pages. It returns long descriptions that can supply substantial context to an agent.

Exa combines keyword methods with embedding-based search, which compares semantic representations rather than relying only on matching terms. Cloudflare configures it to return query-relevant highlights from retrieved pages. That format suits prompts that need concise evidence.

Linkup returns sourced snippets through a fast search mode without generating a synthesized answer. This keeps retrieval separate from reasoning. The application can decide which model analyzes the results and how citations appear.

Those distinctions give Cloudflare a reason to support multiple providers. Search quality is not one-dimensional. Freshness, index coverage, semantic relevance, latency, snippet length, and source selection can each matter more for different tasks.

The same diversity also complicates the product. A normalized response does not make providers equivalent. Developers still need evaluations that measure whether each service retrieves the right evidence for their domain.

Cloudflare therefore pressures search companies in two directions. It gives them distribution through an established developer platform, but it also presents them as interchangeable options behind a common interface. That can shift part of the customer relationship toward the gateway.

Model providers face a different challenge. Their integrated search tools can optimize retrieval and generation together. Cloudflare’s approach argues that many teams will value portability, independent provider choice, and centralized policy more than a tightly coupled experience.

Vercel is pursuing a related strategy. Its AI Gateway recently added search and fetch tools that work across models supporting tool calls. The proximity of these launches suggests that gateways are expanding beyond model routing into complete agent tool layers.

This is not a contest over who first connected a model to search. It is a contest over which platform governs the connection. The winner controls credentials, logs, routing decisions, retention settings, and the developer’s integration surface.

The Mechanism Is Simple, but the Architectural Shift Is Larger

Cloudflare turns search into a reusable retrieval primitive that applications can invoke before, during, or between model calls.

A basic workflow starts with a user question. The application sends that question, or a derived query, to the Web Search API. It receives structured results and places selected descriptions inside a model prompt.

An agent can also decide when to invoke the service. The developer defines a web_search function tool, which is a callable operation described to the model. When the model requests that tool, application code executes the search and returns the results.

The model then receives a second inference call containing the original question and retrieved evidence. This pattern is often called tool-augmented generation because the model gathers external information while completing a task. It differs from relying only on knowledge stored in model parameters.

Cloudflare says native server tools are coming later. Those tools would move more orchestration into AI Gateway itself. For now, developers must implement the loop that receives a tool call, runs a search, and returns the result to the model.

That distinction is important. The current beta offers a search service and common gateway controls. It does not yet provide a fully managed research agent that determines queries, filters evidence, resolves conflicting sources, and composes a cited answer.

The standalone design still has useful advantages. Teams can inspect raw results before they reach the model. They can reject blocked domains, require approved sources, remove duplicate pages, or limit retrieval to documents meeting an internal trust policy.

A support assistant offers a practical example. When a customer asks about a recently changed API, the agent can search current documentation instead of answering from an older training snapshot. The application can preserve retrieved URLs for review.

A software agent can use the same pattern when resolving errors. It might search for a new release note, a changed configuration option, or a current compatibility warning. The search results become evidence, while the model remains responsible for interpreting that evidence.

Research and monitoring products can issue several focused queries. One request might locate an official announcement, another might find documentation, and a third might check independent reporting. A good workflow compares those sources instead of accepting the first result.

Knowledge workers face a related problem inside private material. A system must combine external developments with notes, documents, and established organizational context. That process resembles knowledge blending, where new evidence becomes useful only after connecting with information a user already trusts.

The gateway can log the retrieval calls involved in such workflows. That gives operators a clearer record when an answer fails. They can ask whether the query was poor, the provider missed a page, the snippet lacked context, or the model misread good evidence.

This separation supports better evaluation. Retrieval quality and answer quality can be measured independently. Without that split, a low-quality response reveals little about whether search or generation caused the problem.

It also enables staged fallbacks. An application might query one provider, inspect the result count, and retry with another provider when coverage is weak. Cloudflare has not claimed that the beta automatically makes such quality decisions, so developers must build and test them.

The same applies to caching. Some current-event questions become stale quickly, while stable documentation queries can reuse results. A responsible application needs policies for expiration, source updates, and repeated requests.

Cloudflare’s existing gateway search support already proxies native tools from several model providers. The new API adds a different route: a provider-agnostic retrieval call independent of those native tools.

That means teams now have two search patterns inside the broader Cloudflare platform. They can preserve a model provider’s native tool behavior, or use the new standalone API. The right choice depends on portability, control, and how much orchestration a team wants to own.

The mechanism looks like a modest API addition. Architecturally, it gives search the same status as inference, storage, queues, and other composable services. Agents can treat current web information as infrastructure rather than a special feature bundled with one model.

AI Gateway Web Search Raises the Stakes for Observability

Centralized logs are most valuable when teams can connect a generated claim to the exact retrieval steps that supported it.

Cloudflare positions AI Gateway as the control plane for model applications. It already offers request visibility, security controls, and usage management. Adding retrieval expands that observability into a stage that often determines whether an answer is current.

A model can reason carefully and still produce a wrong response when its evidence is incomplete. Search can return an outdated page, a copied article, or a result that shares the right vocabulary but addresses another subject. Operators need visibility before they can distinguish these failures.

Request identifiers and latency metadata provide a starting point. They can help correlate slow or unsuccessful searches with a particular agent run. Logs can also reveal whether an application is issuing unnecessary queries or repeatedly retrieving the same pages.

However, logging search traffic creates its own governance questions. User queries can expose confidential plans, customer names, security incidents, medical concerns, or internal project details. A centralized gateway must therefore make access boundaries and retention behavior clear.

Cloudflare says all three launch partners support Zero Data Retention for requests routed through this service. Zero Data Retention means a provider does not retain request data after processing under the applicable arrangement. It does not automatically answer every privacy question across the full application.

The developer still controls what enters the query. Cloudflare still operates the gateway and its logs. The final model provider receives whatever retrieved context the application sends. Each stage requires a deliberate data policy.

Bring-your-own-key support also needs careful configuration. Cloudflare’s documentation says an explicit alias causes the request to fail when that key is unavailable. Without an explicit alias, the gateway can use a configured default key or available gateway credits.

That behavior offers convenience, but teams should decide whether fallback is acceptable. A regulated workload may require a specific provider agreement. Silent movement to another commercial path can conflict with internal controls, even when the technical result is valid.

Observability must also preserve enough detail for evaluation without storing excessive sensitive data. Counts and latency alone cannot explain relevance failures. Full queries and result snippets offer more diagnostic value, but they also increase exposure.

The right balance will differ by application. A public news assistant can log more retrieval detail than an internal legal research system. Cloudflare’s advantage depends on whether administrators can express those differences through understandable policies.

Operational control also includes abuse prevention. An agent caught in a tool loop can issue many repeated searches. Rate limits, request budgets, and per-application permissions matter because retrieval can become a significant share of an agent’s activity.

Centralized logs help identify that behavior. They do not prevent it by themselves. Teams still need maximum tool-call counts, timeouts, domain policies, and clear stopping conditions inside their agent framework.

The same principle applies to security. Search results contain untrusted text, and webpages can include instructions aimed at manipulating an agent. Prompt injection occurs when external content tries to override an application’s intended rules.

A normalized search response does not neutralize malicious content. The model can still interpret a hostile snippet as an instruction. Developers should label retrieved material as evidence, restrict the actions available after retrieval, and avoid placing secrets inside tool-enabled contexts.

The new API makes these practices easier to centralize, but it cannot replace them. Cloudflare is selling a better control point. Customers remain responsible for the behavior of the agent operating behind that point.

Crawler Rules Create a Useful Promise and a Hard Test

Cloudflare is tying the launch to responsible crawling, but compliance does not guarantee complete, accurate, or representative search results.

The company requires participating providers to identify their crawlers, respect robots.txt, and include links to retrieved content. It also says provider crawlers must satisfy Cloudflare’s requirements for verified bots.

A verified bot is an automated service whose identity Cloudflare has confirmed. Verification gives site operators a clearer signal when deciding whether to permit or block a crawler. It reduces ambiguity compared with unidentified traffic claiming to represent a search company.

Cloudflare presents this as a standard for a fairer relationship between AI search services and publishers. The policy gives creators more visibility into who accesses their pages. Source links also make it possible for users to inspect the underlying material.

Those commitments distinguish the launch from opaque scraping. They are especially relevant because publishers increasingly question how AI systems acquire, summarize, and commercialize web content. Search providers need access, while site owners want enforceable control.

The tradeoff is that respectful crawling can reduce coverage. Some sites block automated access, restrict particular bots, or place material behind authentication. Search results can only reflect pages a provider indexed and remained permitted to use.

The three providers may therefore return different views of the web. They operate separate indexes, ranking systems, refresh schedules, and snippet-generation processes. A shared API schema hides those implementation details without removing their effects.

Source attribution presents another challenge. Returning a URL is necessary, but it does not prove that a generated answer accurately represents the page. Applications must preserve the connection between each claim and its supporting result.

The final model can combine several snippets into a statement that no source explicitly supports. It can also overlook publication dates or confuse an updated document with an older version. Citation presence should not be mistaken for citation accuracy.

Search ranking adds further uncertainty. Highly optimized pages can outrank primary sources. Syndicated copies can appear more prominently than original reporting. A retrieved description can omit caveats that become obvious on the full page.

Cloudflare’s current API returns search results rather than full-page verification. An application that needs high confidence should fetch important pages, inspect their contents, compare dates, and prefer primary documents. One search call is discovery, not proof.

Latency can also shape quality decisions. Agents often face response-time budgets, so developers may select a fast provider or stop after the first plausible result. That optimization can conflict with the need to verify a sensitive claim through several sources.

The beta status matters here. Cloudflare has documented the interface and provider characteristics, but public evidence about comparative relevance remains limited. Developers should treat provider descriptions as design guidance, not independently verified performance rankings.

There is also no universal benchmark for every application. A provider that performs well on software documentation might struggle with local news, scientific literature, or obscure company filings. Teams need test sets drawn from their own expected queries.

A useful evaluation should record whether the correct page appears, how highly it ranks, how fresh it is, and whether the snippet preserves essential context. It should also test adversarial pages, ambiguous names, and queries with changing answers.

Cost belongs in that evaluation even when exact commercial terms vary. Agentic workflows can multiply one user question into several searches and model calls. Teams should measure total task cost rather than comparing an isolated request.

Cloudflare’s no-markup positioning reduces one concern, but it does not determine value. A more expensive retrieval path can be worthwhile if it avoids additional calls or improves answer accuracy. A cheaper path can become costly when weak results trigger retries.

The crawler standard is still a meaningful part of the announcement. Cloudflare is using its position between sites and automated clients to set participation requirements. The test is whether those rules produce accountable retrieval without creating blind spots that developers fail to notice.

What Developers Should Watch After the Beta Launch

The next phase will show whether Cloudflare Web Search API becomes a durable gateway primitive or remains a convenient wrapper around partner endpoints.

The first signal is the arrival of native server tools. Cloudflare says web search will be among the first tools integrated directly into the AI Gateway control plane. That release would reduce the orchestration code developers currently maintain.

A useful server-tool implementation must do more than hide a function call. Developers should watch how it handles tool permissions, maximum uses, retries, timeouts, result provenance, and provider selection. Those controls determine whether teams can safely use it in production agents.

If server tools preserve common behavior across models, Cloudflare’s gateway thesis becomes stronger. The platform would own both model access and retrieval orchestration. If each model still requires substantial custom handling, the benefit narrows to billing and observability.

The second signal is measurable provider portability. Cloudflare’s shared response format makes switching look simple at the API level. Real portability requires comparable result quality, predictable error handling, and stable behavior under production load.

Teams should run the same query sets through Ceramic.ai, Exa, and Linkup. They should compare source coverage, freshness, ranking, latency, and snippet usefulness. They should also inspect how results change for ambiguous or adversarial questions.

Provider-specific strengths are valuable only when developers can select them deliberately. If most applications remain on the default without evaluation, the marketplace element becomes less meaningful. If teams route by workload, Cloudflare gains a defensible coordination role.

The third signal is competitor response. Vercel already exposes gateway-level search tools, while model companies continue improving native browsing. Other cloud platforms can combine search, models, and agent runtimes inside their own control planes.

Watch whether these competitors add more independent retrieval providers, stronger evaluation tools, or unified citation formats. That reaction will reveal whether provider-agnostic search becomes a standard gateway feature or a temporary point of differentiation.

Developers should also monitor changes to bot policies and publisher controls. Retrieval quality depends on continued access to useful sources. A growing divide between crawlable and restricted content would affect every provider, even when each crawler follows stated rules.

For an initial trial, choose a task with answers that change often and have identifiable primary sources. Release notes, service status, product documentation, and public filings offer clearer evaluation targets than broad opinion questions.

Create a small benchmark before integrating search into a user-facing agent. Record the expected source, acceptable publication date, and facts that a correct answer must include. Then test retrieval and generation separately.

Preserve source URLs in the application interface whenever possible. Users should be able to inspect evidence, especially when an answer affects a consequential decision. A citation should support a specific claim rather than decorate an entire response.

Set a search budget for each task. Limit repeated queries, stop circular tool calls, and require additional confirmation before an agent takes an external action. Search gives a model new information, but it does not give that information authority.

The Introducing Web Search API via AI Gateway launch ultimately asks developers to reconsider where search belongs. Is it a feature of a model, a direct vendor relationship, or a shared service controlled by the application platform?

Cloudflare has made a credible case for the shared-service model. The beta combines three providers, one interface, gateway logs, credential management, and explicit crawler standards. Its value will depend on reliability, retrieval quality, and the promised server-tool layer.

For teams already using AI Gateway or Workers, the practical next step is a controlled evaluation against real queries. Compare the three providers, inspect every source, and measure complete task outcomes. Will an independent search layer improve your agent, or will another gateway boundary create more work than it removes?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page