top of page

Databricks ai_decide Moves Governed Data From Analysis to Action

6 hours ago
12 min read

Databricks launched Databricks ai_decide on September 30, adding a beta SQL function that turns governed data into probabilities, choices, and scores. The conflict is immediate. Enterprises want AI-assisted decisions at data speed, but operational choices demand more accountability than ordinary text generation.

The new function evaluates structured records or text against a rubric supplied by the user. It can estimate whether an event requires attention, choose among named outcomes, or score a case on an ordered scale. Those outputs can then feed another SQL query, workflow, or application.

That makes the announcement more consequential than another model endpoint. Databricks is moving model-based judgment inside the data workflow, where teams already control tables, permissions, pipelines, and business logic. Google Cloud offers related generative functions in BigQuery, while general model APIs let developers construct similar systems manually. The contest now concerns who can make probabilistic decisions operational without hiding their uncertainty.

Databricks ai_decide Turns One SQL Call Into Several Decisions

The important change is not that Databricks can call a model from SQL. It is that one governed function can return several decision-ready assessments about the same record.

According to the company’s launch post, Databricks ai_decide is intended for fast decisions over governed enterprise data. The function sits within the company’s broader family of task-specific AI Functions.

Its syntax has three parts: a state, a collection of questions, and optional version settings. The state contains the evidence being evaluated. It can be ordinary text, a JSON-encoded object, a JSON array, or a VARIANT generated by another AI Function.

Questions are defined once for the call and applied to each input row. Every question includes instructions and a response type. Some types also require criteria describing the available outcomes.

The function reference documents three response types:

  • noul estimates the probability that a statement is true, returning a number between 0 and 1.

  • choice selects one label from as many as 255 named criteria and returns probabilities for every label.

  • score evaluates the input against an ordered scale containing between 2 and 10 criteria.

The unusual term noul identifies a probabilistic yes-or-no assessment. Instead of forcing a Boolean answer, the function reports estimated likelihood. That distinction matters when downstream systems need thresholds rather than absolute claims.

A support organization, for example, could ask whether a ticket needs immediate escalation. It could also ask which team should own the case and how urgent the situation appears. All three assessments can use the same ticket as evidence.

The result is a VARIANT containing a response, metadata, and an error field. A VARIANT is a flexible data type for semi-structured values, such as nested JSON. Successful calls identify the function version, while failed calls can return an error description.

For choice questions, the output includes the selected label, probabilities for each possible label, and a confidence value. Score questions include a numerical score, the original scale descriptions, probabilities, and confidence.

This design gives analysts more information than a single generated label. A workflow can accept high-confidence decisions automatically, route uncertain cases to people, and record the probability distribution for later review.

Databricks also warns that generated answers can vary between calls. That statement is easy to overlook, but it defines the core operational challenge. SQL syntax makes the function accessible and composable. It does not make the underlying judgment deterministic.

Governed AI Decisions Put Data Teams Under Pressure

Databricks ai_decide pressures data teams to treat model judgment as production logic, not an experimental output copied from a chatbot.

Many enterprise decisions already begin in a warehouse or lakehouse. Support tickets, product listings, insurance documents, incident reports, applications, and transaction records eventually become rows processed by pipelines.

Traditional SQL works well when the decision can be written as an exact rule. A transaction above a fixed amount can enter a review queue. A ticket carrying a known error code can go to a specific team.

The harder cases depend on meaning. A customer may describe a service outage without using the company’s official incident vocabulary. A product listing may imply suitability without matching a controlled taxonomy. A case may satisfy several competing priorities at once.

Organizations often handle these situations through manual queues or external model services. Manual review can be slow. External services introduce additional code, data movement, credentials, monitoring, and governance work.

Databricks ai_decide compresses that path. A team can express a qualitative rubric beside the data and receive structured assessments within SQL. Those assessments can then participate in filters, joins, dashboards, Lakeflow pipelines, Workflows, or application logic.

The broader AI Functions overview describes built-in functions for document parsing, extraction, classification, search preparation, and other transformations. Databricks ai_decide adds an explicit decision layer after those preparation steps.

Consider a document workflow. ai_parse_document can convert an uploaded document into structured content. ai_extract can identify specified fields. Databricks ai_decide can then assess the resulting VARIANT against a business rubric.

That sequence changes who can build the workflow. A data engineer no longer needs to wrap every decision in a custom service. An analyst can inspect the state, criteria, answers, probabilities, and errors through familiar data tools.

It also changes who becomes accountable. Once an AI-generated score controls routing or prioritization, the data team owns more than query performance. It must help define acceptable error rates, escalation thresholds, monitoring rules, and fallback behavior.

Governance becomes part of the product design. Databricks says document data remains within its security perimeter. The company says it does not store parameters passed into AI Function calls, although it retains run metadata such as the runtime version.

Access is not automatically narrow. Databricks documentation says users receive EXECUTE permission on the system.ai schema by default when the relevant preview is enabled. Administrators must remove that schema-level permission before granting access to selected functions or groups.

Those access controls are themselves in public preview and require enablement. They apply to task-specific functions under system.ai, but they do not govern the general-purpose ai_query function.

That boundary matters. An enterprise cannot assume that enabling one governance mechanism covers every route to a model. Administrators need separate policies for task-specific AI Functions and direct Model Serving calls.

The launch therefore creates pressure on platform owners, security teams, and operational leaders at the same time. Platform owners must make the function reliable. Security teams must configure access intentionally. Business leaders must define where probabilistic automation is acceptable.

Structured Rubrics Are the Real Mechanism

The central mechanism is a shift from open-ended prompting to explicit rubrics with structured uncertainty.

A general model prompt can ask, “What should we do with this case?” The answer may be articulate, but another system must parse it. The model might also invent a category, change its format, or explain a decision without producing a dependable field.

Databricks ai_decide narrows the interaction. The developer defines named questions, instructions, and allowed criteria. The function returns a predictable response shape that downstream SQL can address.

That constraint reduces integration work. It also exposes the decision contract to reviewers. A compliance specialist can examine the escalation definition. An operations manager can inspect category descriptions. A data engineer can verify how probabilities become workflow actions.

The choice type illustrates the approach. Suppose a support organization defines shipping, billing, and technical support as the only routing labels. The function must select from those names and return a probability for each one.

The selected label is useful, but the distribution can be more informative. A result split closely between billing and technical support signals ambiguity. A workflow can send that case to a general queue instead of pretending the top label is certain.

The score type uses ordered criteria rather than a free-form number. A team might define three urgency levels, from a routine request to a critical blocker. The returned score is a probability-weighted average of the criterion indexes.

That method preserves information about competing assessments. If the model assigns probability across several levels, the result can be fractional. The output also retains the legend and the probabilities behind the score.

Multiple questions can share one state. That reduces the need to pass the same evidence through separate prompts for category, urgency, escalation, and other judgments. It also keeps the related answers together.

However, sharing input does not guarantee that every question represents an independent assessment. Teams should test whether instructions interact in unexpected ways. They should also verify whether combining questions changes quality, latency, or cost for their workload.

Databricks says the underlying model can change if another model performs better in its internal benchmarks. The current documentation associates possible models with the Apache 2.0 license and directs customers to applicable model terms.

A managed model choice reduces configuration. It also means the function’s behavior can evolve beneath a stable SQL interface. Version metadata helps identify the function contract, but teams still need regression tests based on representative data.

That is where the mechanism becomes operationally important. A stored procedure built from deterministic conditions can be tested against exact expected outputs. A probabilistic function needs distribution checks, threshold checks, and repeated evaluations.

Teams should maintain labeled evaluation sets for each important rubric. Those sets should include common examples, borderline cases, missing evidence, conflicting evidence, and inputs that should always reach a person.

They should also separate recommendation from execution. A decision function can prioritize a support queue with limited downside. The same confidence should not automatically authorize a refund, reject an applicant, suspend an account, or initiate a safety response.

The SQL interface makes composition easy. Good system design must keep consequential actions deliberately difficult.

Databricks ai_decide Competes With General Model Calls and Warehouse AI

The main contest is between a task-specific managed function and the flexibility of building decision logic around a general model endpoint.

Databricks already offers ai_query, a general-purpose function that invokes a Model Serving endpoint. Developers can choose a supported model, write their own prompts, and control parameters and return types.

The ai_query documentation recommends that teams begin with a task-specific AI Function when one matches their objective. It positions ai_query for cases requiring more control over the model, prompt, parameters, or output.

That distinction creates a clear tradeoff.

A task-specific function reduces setup and imposes a structured contract. Databricks manages the system behind the operation and can improve its implementation. Teams can focus on their evidence, questions, and criteria.

A general model call offers flexibility. Developers can use a custom model, tune decoding settings, define a different schema, implement fallback endpoints, or preserve a fixed model version. They also assume more engineering and evaluation work.

Databricks ai_decide is strongest when a decision fits its three available forms. Probability, named choice, and ordered score cover many routing and prioritization tasks. They do not cover every decision structure.

A business may need multilabel classification, constrained numerical estimates, citations to evidence, rule-based exclusions, or a chain of dependent questions. Developers may still need ai_query, custom functions, or an external application for those cases.

The competitive field also extends beyond Databricks. Google Cloud documents an AI.GENERATE_BOOL function for BigQuery that returns a Boolean result, response details, and status information. It can process text and referenced unstructured content through Gemini.

Google’s boolean function supports model and request parameters. Its documentation also warns that prompt design affects results and that query planning can make model inference process more rows than expected.

Google separately offers AI.IF, which its documentation describes as supporting prompt optimization and an optimized mode. That mode can train a distilled model for lower cost and latency at scale.

Databricks takes a broader rubric-oriented approach in one call. Databricks ai_decide can answer multiple questions and return probabilities for named choices or ordered scores. Google’s documented Boolean function focuses on true-or-false generation, although BigQuery offers additional scalar and generative functions.

Neither approach eliminates application design. Warehouse-native AI reduces the distance between data and inference, but teams still choose thresholds, materialize inputs, control permissions, and evaluate outputs.

Competition will therefore turn on operational evidence rather than syntax alone. Buyers need to know how functions behave on their records, in their regions, under their compliance requirements, and at production volume.

A function that saves integration work but creates unpredictable costs will struggle. So will a flexible endpoint that demands a specialist team for every routine classification task.

Databricks is betting that many enterprise decisions share enough structure to deserve a managed primitive. The outcome depends on whether those primitives remain understandable when organizations connect them to real actions.

Fast Decisions Still Need Slow Validation

The beta label is the clearest warning: Databricks has simplified implementation, but it has not removed uncertainty, regional limits, or human accountability.

Databricks ai_decide is available as a beta feature. Workspace administrators control access through the Previews page, and the function is available only in supported regions.

It does not run on Databricks SQL Classic. The documentation requires Databricks Runtime 15.4 LTS or later and recommends Runtime 18.2 or later for current features and performance.

Those prerequisites narrow immediate adoption. Organizations with older runtimes, Classic warehouses, unsupported regions, or strict preview policies must change infrastructure or wait.

The model layer introduces another uncertainty. Databricks says it might change the underlying model when its internal benchmarks identify a better option. That managed evolution can improve results, but it also creates model drift from the customer’s perspective.

A decision pipeline cannot rely solely on the function name remaining stable. Teams need baseline evaluations, release controls, monitored thresholds, and the ability to compare results after platform changes.

The rubric itself can also fail. Instructions may omit a meaningful exception. Categories may overlap. An ordered scale may suggest precision that the evidence does not support.

Confidence requires careful interpretation. A high confidence field indicates how well the state supports the assessment under the function’s process. It does not establish that the answer is factually correct or fair.

Probability outputs carry a similar risk. A value of 0.9 looks precise, but users should not assume that 90 percent of comparable predictions will be correct without calibration evidence. Calibration must be tested on the organization’s own labeled cases.

Data quality remains decisive. If the state contains stale, incomplete, or misleading records, a well-structured rubric can still produce a poor decision. Governance controls who may use data; they do not guarantee that every input is suitable.

Bias can also enter through examples and criteria. A category description may encode a team’s historical practice, including its blind spots. A score may reproduce inconsistent human judgments from an evaluation set.

High-impact uses need stronger safeguards. Employment, credit, healthcare, insurance, legal, and safety decisions involve obligations that extend beyond a model output. Organizations should involve domain, legal, security, and risk teams before automating actions.

Even lower-risk workflows need failure handling. The function can return a null response and an error message. Pipelines must decide whether to retry, pause, use deterministic fallback logic, or route the case to a person.

Generated answers can vary between calls, according to the documentation. Repeated evaluations may therefore produce different labels or scores for borderline records. Teams need idempotency policies when downstream actions should occur only once.

Cost deserves similar scrutiny. Each model-backed assessment consumes inference resources. Running several questions over every row of a large table can turn a convenient query into an expensive operation.

Developers should isolate the relevant rows before invoking the function. They should materialize stable input sets, prevent accidental full-table reevaluation, and store outputs when repeated inference adds no value.

None of these concerns invalidates the product direction. They explain why governed AI decisions need a different standard than generated summaries. A weak summary inconveniences a reader. A weak routing or prioritization decision changes what happens next.

Three Signals Will Show Whether the Bet Works

The next test is whether Databricks can turn a promising SQL abstraction into a measurable, governable production capability.

The first signal is documented evaluation performance. Databricks has not presented independent benchmarks establishing accuracy, calibration, latency, or cost across representative decision tasks. Customers need workload-specific evidence rather than a general promise of speed.

Useful validation would compare Databricks ai_decide with ai_query, deterministic rules, and established human review. It should measure label accuracy, probability calibration, stability across repeated calls, processing time, and the number of cases requiring escalation.

If teams can publish repeatable gains with controlled error rates, the task-specific approach gains credibility. If they must wrap every call in extensive corrective logic, the abstraction saves less work than the syntax suggests.

The second signal is wider production availability. The feature currently carries a beta label, requires preview enablement, excludes Classic SQL warehouses, and supports only certain regions.

Movement toward broader availability would indicate that Databricks is confident in service reliability, governance coverage, and operational support. Persistent preview restrictions would limit the function to experiments and low-risk workflows.

Governance maturity belongs in the same signal. Administrators need clear controls for execution permissions, model access, audit trails, regional processing, and changes to underlying models. These controls must work consistently across cloud platforms and workspace configurations.

The third signal is customer use beyond demonstrations. Product categorization and support-ticket routing are understandable examples. The stronger evidence will come from production workflows with published review thresholds and measurable business outcomes.

Watch for cases where probabilities change the workflow, not merely decorate a dashboard. A company might automate clear decisions, send ambiguous records to specialists, and use the resulting corrections to test its rubric.

Also watch competitors. Google Cloud already exposes warehouse-native generative functions, and other data platforms continue adding model access near governed data. A competing function with clearer calibration, broader input support, or lower operational overhead would weaken Databricks’ advantage.

Databricks ai_decide captures an important shift. Enterprises are no longer satisfied with models that only describe information. They want systems that help choose, rank, route, and escalate while remaining inside established data controls.

The prudent next step is not to connect the function directly to a consequential action. Select one bounded queue, define an explicit rubric, and build a labeled evaluation set. Compare the function’s outputs with current decisions, then choose thresholds for automation and human review. Track errors, uncertainty, latency, and drift separately.

Which decision in your organization is repetitive enough to evaluate, yet reversible enough to test safely? That is the right starting point for Databricks ai_decide. The product’s lasting value will come from transparent operating discipline, not from treating probabilistic output as certainty.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page