top of page

Imp DSPy Port Brings Optimizable AI Programs to the BEAM, but Production Proof Comes Next

2 hours ago
13 min read

Imp has released an Imp DSPy port for the BEAM with one ambitious promise: bring optimizable language-model programs into Elixir’s process-oriented runtime. The first Hex release includes typed signatures, reasoning modules, evaluation, optimizers, retrieval, and supervised agent runs. That breadth makes Imp more than another wrapper around a model API.

The project calls itself a full port of DSPy, the Python framework for building language-model programs that can be measured and optimized. Imp retains that programming model while changing the host environment. An Imp program is an immutable Elixir value, and an agent can run as a supervised BEAM process.

That combination creates the real tension. Python remains the center of AI framework development, while Elixir excels at concurrent, long-running services. Imp argues that developers should not have to choose between DSPy-style optimization and the operational model of Erlang/OTP.

The code is available now, but the production verdict is not. Imp 0.5 is experimental, its API can change, and its optimizers still need broader benchmarking. The release therefore establishes technical scope, not proven parity across every workload.

The Imp DSPy Port Reaches Beyond Basic Model Calls

Imp recreates the main DSPy programming model instead of translating only its simplest prediction interface.

The Imp repository describes the project as a complete port of DSPy to the BEAM. Its public surface covers signatures, modules, examples, metrics, evaluation, optimizers, tools, retrieval, and saved programs. It also includes agent loops and process-based execution.

A signature is a typed declaration of what a model step receives and returns. Developers describe a task such as an issue, classification, and summary without manually assembling every prompt. Imp then formats the request, calls the selected model, parses the response, and validates its fields.

That structure follows the central idea behind DSPy programs. DSPy treats model behavior as a program that can be evaluated and improved, rather than a collection of handwritten prompt strings. Imp carries that idea into Elixir while preserving familiar names and concepts.

The basic example in Imp defines a GitHub issue triage task. Its output restricts the issue type to bug, feature, or question, alongside a generated summary. If the model returns an invalid type, the call produces an error instead of silently passing malformed data downstream.

Developers can replace a direct prediction with chain-of-thought reasoning or a ReAct agent without changing the signature. ReAct is a loop in which a model selects tools, observes their results, and continues until it returns an answer. The task contract remains separate from the reasoning strategy.

Imp also exposes several DSPy-style optimizers. LabeledFewShot selects examples, BootstrapFewShot generates additional demonstrations, and MIPROv2 searches across instructions and examples. SIMBA learns from stronger and weaker attempts, while GEPA reflects on failures and proposes revised instructions.

These components matter because optimization is the feature that separates DSPy from ordinary model client libraries. A client library standardizes requests. An optimizer repeatedly evaluates program variants against a metric and returns the strongest observed configuration.

Imp asks developers to divide examples into training, validation, and test sets. A metric scores the program, while an optimizer changes instructions, demonstrations, or related parameters. The resulting program can be inspected, saved as JSON, and compared with its earlier version.

The project also supports retrieval, best-of-N selection, output refinement, program-of-thought execution, and recursive language-model workflows. It includes MCP tool imports and ACP serving, connecting Imp programs with external tools and compatible agent hosts.

This is a wide initial surface. It supports the claim that Imp targets DSPy’s architecture, not merely its terminology. However, feature presence does not settle behavioral parity, performance, or operational maturity.

The release documentation acknowledges that distinction. Imp 0.5 is the project’s first Hex release, and the maintainers describe it as experimental. They also warn that its API may change and that large-scale optimizer benchmarking remains unfinished.

That warning is central to interpreting the launch. Imp has delivered a substantial implementation, but “full port” is still a project claim. Independent testing must show how consistently its modules and optimizers match DSPy under realistic conditions.

Why the BEAM Changes the Agent Runtime

The important change is not Elixir syntax; it is the ability to model each long-running agent as an isolated, supervised process.

The BEAM is the virtual machine used by Erlang and Elixir. It schedules many lightweight processes that communicate through messages and maintain isolated state. OTP adds established patterns for supervision, fault handling, and long-running services.

Imp uses those properties directly. A normal call can execute inside the caller’s process, while start_run launches a program as its own supervised process. The caller can monitor that run, stop it, collect events, and control which tool calls receive authorization.

Elixir’s GenServer model shows why that approach is different from adding asynchronous functions to a Python library. A GenServer is a process that keeps state, handles synchronous and asynchronous messages, and fits into a supervision tree.

For an AI agent, that model creates a natural home for state and lifecycle management. One process can represent one agent run. Other processes can monitor it, receive events, impose deadlines, or restart surrounding services without sharing mutable memory.

Imp records events such as run creation, model requests, model responses, tool calls, tool results, and completion. Those events create an observable execution history. They also give optimizers material for evaluating an entire agent trajectory rather than only its final answer.

Tool authorization becomes part of the runtime boundary. The project’s example allows an agent to fetch content only from an approved host. A denied tool call never receives permission merely because the model requested it.

This does not make model-generated actions safe by default. It does make the authorization decision explicit and programmable. That boundary is useful when an agent can read internal systems, execute utilities, or call external services.

Imp also treats uncertain tool outcomes carefully. A timed-out tool may have completed an external action even when the caller never received confirmation. The project reports such outcomes as unknown instead of retrying them automatically.

That distinction addresses a common agent reliability problem. Repeating a read is usually harmless, but repeating a payment, message, deployment, or deletion can create damage. A runtime should distinguish a failed observation from a confirmed failed action.

Deadlines provide another boundary. Imp says model requests and tool execution can be constrained by a deadline attached to the run. When the owner process ends, supervised work can end with it rather than becoming abandoned background activity.

The BEAM also offers concurrency without requiring every application team to invent a new agent scheduler. Multiple processes can run independently, send messages, and fail in isolation. Supervisors define how related processes respond when one component exits.

That design is especially relevant for applications where agents remain active longer than a single web request. Examples include monitoring agents, support workflows, background research jobs, and systems that wait for human authorization.

Python can support all of those workloads. The difference is that Python frameworks usually assemble lifecycle behavior from task queues, asynchronous runtimes, worker systems, and application-specific state management. The BEAM places these concepts near the center of its programming model.

Imp therefore pressures a particular assumption, not the entire Python AI ecosystem. It challenges the idea that DSPy-style programs must remain tied to Python when their production host is a concurrent service.

For Elixir teams, this reduces a language boundary. They can keep model logic, application state, supervision, and surrounding business rules in one runtime. They may avoid operating a separate Python service solely to gain declarative model programming.

The potential value is clearest inside existing Elixir systems. A team running Phoenix, Broadway, Oban, or other BEAM workloads can integrate an Imp program using familiar deployment and observability patterns. The new component becomes part of the application rather than an adjacent AI island.

That architectural fit is the release’s strongest argument. Syntax parity can be copied. A runtime model built around process isolation, message passing, and supervision changes how developers can operate agents after deployment.

Imp Versus DSPy Is a Host-Runtime Choice

The primary contest is not Imp against DSPy as competing products; it is BEAM-native operation against Python-centered AI development.

DSPy remains the reference point. Its ecosystem, research history, documentation, contributor base, and production examples give it an advantage that a first Hex release cannot reproduce immediately. Imp inherits ideas from that work but does not inherit its accumulated validation.

The project’s DSPy mapping makes the relationship explicit. DSPy signatures map to Imp signatures, Predict maps to Imp.predict, and ReAct maps to Imp.react. Evaluation, retrieval, parallel execution, saving, and several optimizers have corresponding interfaces.

Imp says it tracks DSPy 3.3.1 as of September 2026, while work on DSPy 3.4 additions continues. That detail shows both the project’s ambition and the maintenance burden ahead. DSPy can evolve faster than a separate implementation can follow.

A port must decide where exact compatibility matters and where the host language should shape the design. Imp does not attempt to make Elixir look exactly like Python. Programs are immutable values, model dependencies can be passed explicitly, and context is scoped to the calling process.

That is a sensible approach because direct source compatibility is not the goal. An Elixir developer cannot copy a Python application unchanged. The useful target is conceptual and behavioral compatibility across signatures, modules, metrics, optimizers, and saved artifacts.

Imp’s maintainers have built differential checks against pinned DSPy versions. The repository includes parity gates for prompt templates and tests intended to compare behavior. Its build configuration references a pinned DSPy 3.2.1 environment for golden-trace comparisons.

Those checks are meaningful evidence of engineering intent. They show that the project is measuring compatibility instead of relying entirely on similar method names. Still, repository tests are not independent benchmarks.

The hardest parity questions involve optimizers. Prediction modules can be compared using known inputs and outputs. Optimizers include randomness, repeated model calls, search strategies, budgets, and dataset-dependent behavior.

Imp’s implementation of GEPA illustrates the difficulty. GEPA is an optimizer that reads execution traces, reflects on failures, and proposes new instructions. Imp includes DSPy-oriented execution profiles and a separate BEAM-native profile with different options.

According to Imp’s changelog, its default DSPy profile pins behaviors such as random-number generation, budgets, merge settings, and selection rules. These details can materially affect which program an optimizer returns.

Imp also extends optimization into supervised agent runs. GEPA can inspect thoughts, tool calls, tool results, and final outputs from a trajectory. That feature aligns the optimizer with Imp’s process-based runtime rather than treating agents as opaque calls.

DSPy itself continues to move. Its optimizer catalog includes several strategies for demonstrations, instructions, fine-tuning, and combined optimization. Keeping pace requires more than implementing a fixed API once.

This maintenance race is the central cost of a full port. Every new DSPy module, adapter, optimizer, or behavioral change creates a decision for Imp. The project must port it, document a divergence, or leave the compatibility claim temporarily behind.

The BEAM side brings its own constraints. Imp 0.5 requires Elixir 1.19 or later and a C and C++ compiler. Two dependencies include native build requirements, and the first compilation needs network access for part of that toolchain.

Those requirements are manageable, but they complicate the story that a BEAM-native package automatically means a simpler deployment. Teams must examine native dependencies, release configuration, protocol adapters, and model-provider connections.

Imp reaches model providers through ReqLLM, an Elixir library that standardizes language-model requests. That provides useful separation between the program framework and provider transport. It also makes ReqLLM compatibility part of Imp’s effective provider coverage.

Choosing between the frameworks therefore depends on system boundaries. A Python-first research team gains little from moving to Elixir only for process supervision. An Elixir product team may gain substantially by avoiding a separate Python service.

The decision also depends on who owns optimization. Data scientists may prefer DSPy’s Python environment and surrounding evaluation tools. Backend engineers may prefer an Imp program deployed beside the services and data flows they already operate.

Imp does not need to replace DSPy to matter. It needs to make DSPy’s programming model credible inside production systems where the BEAM already provides the operational foundation.

The Full-Port Claim Still Needs Independent Tests

Imp’s broad feature list is real, but maturity depends on optimizer quality, behavioral parity, and failure handling under sustained workloads.

The first uncertainty is the meaning of “full.” Imp covers the recognizable DSPy layers, yet its own documentation says it tracks an earlier DSPy version while newer additions are still arriving. Full coverage is therefore a moving target.

Some modules also carry different implementation constraints. Program-of-thought, CodeAct, and recursive language-model features run model-written code through Imp’s restricted interpreter. Their behavior will not necessarily match DSPy’s Python execution environment in every case.

That divergence can be beneficial. A restricted interpreter may offer a narrower and more controllable surface. It can also prevent programs from using libraries or runtime behavior that DSPy users expect.

Saved-program compatibility deserves similar scrutiny. Imp can save programs as JSON, but shared concepts do not guarantee that DSPy and Imp can exchange every artifact directly. Field formats, provider configuration, module state, and optimizer metadata can differ.

Provider behavior is another variable. Two frameworks can generate equivalent prompts yet receive different results because their adapters format messages, tool calls, or structured output constraints differently. Small formatting changes can alter model behavior.

Imp has invested in adapter fidelity. Its changelog describes changes that bring structured values, ReActV2 messages, and GEPA reflection prompts closer to DSPy behavior. That work also reveals how many subtle decisions parity requires.

Every provider adds further edge cases. Streaming responses, parallel tool calls, partial text, usage records, timeouts, and malformed structured outputs vary across APIs. A framework must normalize them without hiding meaningful failures.

The current changelog documents fixes involving streamed tool calls, missing model records, caller cancellation, optimizer instructions, and uncertain tool outcomes. These are normal issues for an early project, but they show where production complexity accumulates.

Large-scale optimizer benchmarking is the most important missing evidence. Imp’s maintainers explicitly say this work remains necessary. Users need comparative results across datasets, models, budgets, and repeated runs.

A useful test should ask more than whether both frameworks finish. It should compare baseline scores, optimized scores, total model calls, token usage, elapsed time, reproducibility, and failure rates. Agent benchmarks should also measure tool accuracy and incomplete actions.

The benchmark should separate framework quality from model variance. Both implementations need the same model, datasets, evaluation metric, budget, and comparable random seeds. Multiple runs are necessary because optimizer searches can produce different results.

Operational tests should measure supervision under failure. Researchers should kill owner processes, interrupt model requests, time out tools, overload queues, and restart surrounding applications. The expected outcome must be explicit for each case.

Security testing matters because agent tools cross application boundaries. Imp offers authorization hooks, but application developers still define the policy. Weak host checks, excessive tool permissions, and unsafe arguments can undermine the runtime boundary.

Long-running state also raises questions. Developers need to know what survives a process restart, how checkpoints are persisted, and how upgraded code interacts with saved programs. Supervision restarts a process, but it does not automatically reconstruct correct business state.

Observability must extend beyond event capture. Teams need searchable traces, cost records, model metadata, tool outcomes, and links between an agent run and the surrounding request. The raw event stream is the foundation, not the finished monitoring system.

Adoption presents another risk. Elixir has an active community, but the AI tooling market remains concentrated around Python and JavaScript. Imp must attract contributors who understand both language-model optimization and BEAM application design.

Documentation will influence that adoption. The project already provides a getting-started path, DSPy migration guide, tutorials, production notes, and Livebook notebooks. Maintaining those materials alongside fast-moving code will require sustained effort.

Version stability matters just as much. Teams will hesitate to place core workflows on an API that may change frequently. A clear compatibility policy and migration path would make the experimental label easier to manage.

None of these concerns negates the release. They define the distance between an impressive implementation and a dependable platform. Imp has made the first part visible; users and contributors now need to test the second.

Developers considering an early evaluation should isolate the experiment. A bounded classification or extraction workflow offers a better starting point than an autonomous agent with broad permissions. It produces measurable outputs and limits operational risk.

Teams should also retain a baseline implementation. Running the same dataset through DSPy and Imp creates direct evidence about quality, latency, and cost. The comparison should use held-out examples that were not visible during optimization.

For production trials, the engineering workflow around findings matters too. Teams need a searchable record of test cases, failures, configuration changes, and benchmark results. Otherwise, promising demonstrations can become unsupported architectural decisions.

Three Signals Will Decide Whether Imp Holds Up

Imp’s next phase will be determined by comparative benchmarks, production adoption, and its ability to track DSPy without losing BEAM-native advantages.

The first signal is a reproducible parity benchmark. Imp already contains differential testing and benchmark infrastructure, but external users need published results they can rerun. The strongest evidence would compare Imp and DSPy across identical tasks and budgets.

Such results should include straightforward prediction, structured extraction, retrieval, tool use, and multi-step agents. Optimizer comparisons should cover GEPA, MIPROv2, and few-shot methods because those features support the port’s main value proposition.

If Imp produces comparable quality and cost across repeated runs, the full-port claim becomes stronger. If results vary materially, users need documentation that explains whether adapters, search behavior, randomization, or runtime differences caused the gap.

The second signal is production use inside real Elixir applications. A credible deployment would show more than an agent answering a question. It should demonstrate supervision, backpressure, tracing, authorization, persistence, upgrades, and recovery from partial tool failures.

Evidence from Phoenix services, job-processing systems, or event-driven applications would be particularly informative. These environments expose the reasons to choose the BEAM. They can show whether process isolation simplifies operations or merely moves complexity.

Case studies should disclose workload shape and failure boundaries. A short-lived extraction endpoint tests different properties from an agent that remains active for hours. Both are useful, but they support different claims.

If Elixir teams report simpler deployment and clearer lifecycle control, Imp’s runtime argument gains weight. If most adopters use only synchronous prediction, the broader agent-process design will remain largely theoretical.

The third signal is how quickly Imp follows DSPy 3.4 and later releases. The project says those additions are being brought in. Update speed will reveal whether a full port is sustainable or whether compatibility gaps accumulate.

Exact feature matching should not be the only goal. Imp should preserve the places where the BEAM changes the design for good reasons. Process-scoped context, supervised runs, explicit dependencies, and careful cancellation behavior can justify deliberate differences.

The maintainers will need a clear compatibility vocabulary. Features could be labeled equivalent, adapted, experimental, or intentionally unsupported. That would make “full port” easier to evaluate without expecting byte-for-byte identity.

Users should also watch Hex release cadence and migration quality. Frequent releases can indicate active development, but repeated breaking changes raise adoption costs. Upgrade guides and stable core interfaces can balance those pressures.

Community activity provides a secondary signal. Issues that receive detailed responses, external pull requests, and independent examples show whether the project is expanding beyond its original author. Contributor diversity matters for a framework with such a wide surface.

Security and dependency maintenance deserve attention as well. MCP connections, native dependencies, restricted code execution, and provider integrations expand the attack surface. Clear advisories and timely fixes will be essential for production confidence.

The decisive question is whether Imp becomes the default way to build measurable AI programs inside Elixir applications. That outcome does not require dominance across the entire AI market. It requires trust among teams already committed to the BEAM.

The Imp DSPy port has made a credible opening move. It presents a surprisingly complete programming surface and connects it to an operating model suited to concurrent services. Its own experimental warning keeps that achievement in the proper frame.

Developers can now test the thesis rather than debate it abstractly. Choose one measurable workflow, create fixed training and test sets, and run the same task through Imp and DSPy. Record quality, model usage, latency, failures, and operational effort.

Then test the part Python comparisons often miss. Run the Imp workflow as a supervised process, interrupt it, deny a tool, and inspect the resulting events. If that lifecycle becomes easier to reason about, the BEAM port has delivered something more important than syntax parity.

The next few releases should show whether Imp can maintain that advantage while matching DSPy’s fast-moving capabilities. For now, the project is best understood as a serious experimental runtime, not a finished replacement.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page