
A production agent can call a model yet fail when a retry repeats a write, an approval cannot resume, or nobody can explain a tool call. Choose the runtime around those failures. This guide compares five agentic AI frameworks against one small task and shows when extra machinery is unnecessary.
This documentation-based comparison was checked on September 23, 2026. It makes no matched performance claim and concerns code frameworks for application teams.
What Counts as an Agentic AI Framework
An agent development framework supplies code-level primitives for a model-driven loop: call a model, expose tools, route the next step, retain enough state, and return a result or interruption. It may also provide tracing, evaluations, and deployment adapters. A hosted console may package those primitives, but a console alone does not tell you how your application handles a failed tool call.
Your application owns authorization, side-effect safety, data retention, and release policy. A framework may surface an approval request; your service must validate the approver and action. One model call followed by a deterministic function may need no framework.
The Production Requirements to Define First
Use the same hypothetical task throughout the shortlist: read a support ticket, look up an order through a read-only tool, draft a proposed refund decision, and pause before any refund write. The example is a design exercise; none of the five implementations below was run for this article.
Before selecting a library, write down six acceptance conditions:
- State: A worker restart during approval must not lose the ticket ID, order evidence, or pending decision. Resuming must not issue the refund twice.
- Tools: The order lookup and refund command have distinct identities and scopes. The refund endpoint rechecks authorization and an idempotency key.
- Human control: A named reviewer sees the proposed amount and reason; rejection ends the action path.
- Testing: A fixed order fixture covers an eligible case, an ineligible case, a denied approval, and a restart at the pause.
- Observability: Each run links its ticket ID, framework version, tool calls, approval event, and outcome without leaking customer data into traces.
- Deployment: The chosen state store and worker model can run in the team's existing environment, with a documented rollback path.
These conditions matter more than agent count. The refund API remains the final authority on money movement.
A Shortlist of Framework Approaches
The shortlist covers LangGraph, Google Agent Development Kit (ADK), CrewAI, OpenAI Agents SDK, and Pydantic AI. The groupings describe a likely starting approach, not exclusive feature labels.
Graph and State-Machine Frameworks
LangGraph makes nodes, edges, and state explicit. Its checkpoint persistence and interrupts fit the refund task when the review pause and resume are central to the design. You can model lookup, proposal, approval, and command as separate steps. The cost is that the team must define state shape, node behavior, checkpoint storage, and upgrade discipline. That is worthwhile for long-running paths with branching and recovery; it is ceremony for a one-turn assistant. LangGraph has Python and JavaScript/TypeScript implementations.

Google ADK now also has a graph-based workflow runtime in its 2.0 line, alongside agent and tool primitives. Google's ADK 2.0 migration notes explicitly flag incompatibilities when moving from Python 1.x, which makes an upgrade rehearsal part of selection. It is a candidate when the team wants structured workflows, sessions, evaluation, and multiple language SDKs under one project. Check the exact feature and deployment route in the language you will use; a shared project name does not imply identical behavior across SDKs.

Role-Based Multi-Agent Frameworks
CrewAI starts naturally with agents assigned roles and tasks, but its Flows add event-driven steps and shared state. For the refund example, a crew of “researcher, analyst, approver” is less important than a Flow that prevents the refund command from running before an authorized human decision. CrewAI also documents checkpointing and human feedback. Choose it when specialist delegation actually improves the task decomposition; adding roles to a linear lookup and draft only adds handoffs to inspect.

Lightweight Tool-Calling Libraries
OpenAI Agents SDK offers agents, function tools, handoffs, guardrails, sessions, approvals, and tracing of model, tool, and handoff events with comparatively few application concepts. Python and JavaScript/TypeScript SDKs exist. It fits a small service that wants a managed agent loop while keeping business routing in application code. Its session history and resumable run state are distinct concerns: decide which one stores the approval boundary, how it is restored, and how a side-effecting tool is protected from replay.
Pydantic AI emphasizes typed inputs, tools, dependencies, and structured outputs; its durable execution integrations can support long waits, while its core remains usable for a small Python agent. It suits a Python team that wants type-checked application seams and existing observability infrastructure. Durability needs a selected engine and storage plan; a typed tool signature by itself does not make a refund command safe to retry.
Compare State, Tools, and Human Control
Compare these documented levers in a forced-restart pilot. The reference surface is each framework's Python documentation, including ADK 2.0. Other language SDKs and deployments may differ.
| Framework | State | Tools | Human stop | Testing | Observability | Deployment |
|---|---|---|---|---|---|---|
| LangGraph | Checkpoints | Node adapters | Interrupts | Node and resume cases | Trace integration | App and checkpoint store |
| Google ADK | Sessions, events | Functions, MCP | Tool confirmation | Eval cases | Events and traces | App or Google runtime |
| CrewAI | Flow state, checkpoints | Crew tools, MCP | Flow feedback | Flow cases | Tracing | App or AMP |
| OpenAI Agents SDK | Sessions, RunState | Functions, MCP | Tool approval | Scripted model | Built-in traces | App service |
| Pydantic AI | Conversation store, durable engine | Typed tools, MCP | Deferred approval | Test model, evals | OpenTelemetry | App and chosen engine |
These are different persistence contracts, not interchangeable memory. Check which state survives after lookup and who can resume it.
Next, wrap the same two tools in each candidate. Keep a framework-neutral application interface such as get_order(ticket_id, order_id) and request_refund(order_id, amount, idempotency_key). The framework adapter should expose only the allowed operation; the service behind it must enforce resource authorization, amount rules, and idempotency. A model-visible tool description is useful guidance, not access control.
Finally, make the approval deny path observable. Can the runtime pause without holding an open process? Can the reviewer see the exact command arguments? Can a rejected request resume into a different write? In the OpenAI SDK, tool approval and guardrail behavior have specific execution boundaries; for any framework, inspect those boundaries instead of treating “human in the loop” as a blanket guarantee.

Compare Testing, Observability, and Deployment
Run the same fixture against each candidate. Stub model responses for the four control cases; evaluate real model behavior separately. OpenAI's deterministic SDK testing utilities and Google's ADK evaluation workflow offer starting points. Neither replaces a refund-service integration test.
Inspect a failed run before inspecting a successful demo. The trace should show the versioned code path, tool arguments after redaction, retry count, reviewer decision, and final external transaction ID. Confirm where traces travel and who can read them. Built-in tracing, an OpenTelemetry export, and a hosted dashboard have different data and retention contracts.
A local checkpoint database may fail with multiple workers. Decide whether you need a shared store, background jobs, an approval webhook, or a managed service. CrewAI's production architecture guidance separates persistence and deployment. Open-source licensing does not include a hosted platform's SLA or data terms.
Choose the Smallest Stack That Fits
Pick LangGraph for explicit transitions and interrupts; ADK when its language SDK and workflow runtime fit your stack; CrewAI when specialist handoffs justify a crew; OpenAI Agents SDK for a small tool-driven service with built-in tracing; Pydantic AI for typed Python boundaries and chosen durability. These are documented feature matches, not a reliability ranking.
For a first production agent, build the refund example with one candidate and reject it if the approval cannot survive restart, the command can be replayed, or the final run cannot be reconstructed. Verdent can help a team plan, implement, and review the repository change, but the application's runtime and the human release owner remain responsible for these controls.
FAQ
Which Agentic AI Frameworks Support More Than One Programming Language?
LangGraph has Python and JavaScript/TypeScript implementations; OpenAI Agents SDK has Python and JavaScript/TypeScript SDKs. Google ADK documents Python, TypeScript, Go, Java, and Kotlin surfaces, though version and feature parity vary. CrewAI and Pydantic AI are primarily Python choices. Verify the particular capability in the intended SDK before treating “multilanguage” as one shared runtime guarantee. LangGraph's language paths and the ADK documentation are useful starting points.
Can an Agent Framework Be Migrated Without Rewriting Every Tool?
Often, but the cutover hinges on runs already in progress. Inventory paused approvals and persisted snapshots, then decide whether the old runtime will drain them or a tested converter can resume them in the new one. Route new tickets to the replacement while old tickets remain pinned to their original runtime; keep a rollback route until the approval, rejection, and retry paths pass on restored snapshots. Shared tool contracts reduce tool rewrites, but they do not migrate serialized state or make an in-flight refund safe to replay.
Which Framework Licenses Permit Commercial Products?
At this check, the open-source repositories for LangGraph, CrewAI, OpenAI Agents SDK, and Pydantic AI use MIT-style licenses; Google ADK Python uses Apache 2.0. Both license families generally permit commercial use under their conditions. Check the exact package, dependencies, notices, trademarks, and any hosted-service terms with your organization's license process. This is a repository snapshot, not legal advice.
Do Agentic AI Frameworks Support MCP Servers Natively?
Several do, including OpenAI Agents SDK, Google ADK, CrewAI, and Pydantic AI. Transport, filtering, approvals, and server lifecycle differ. Verify the exact SDK release and transport you need, then treat the connected server as a separately authenticated and reviewed service.
How Do Framework Upgrades Affect Serialized Agent State?
An upgrade can change a state schema, tool name, node path, or resume behavior while old runs remain paused. Inventory those runs; pin the current version; restore saved snapshots in a staging copy; and test approval, rejection, and retry before rollout. Keep a drain or migration plan if a snapshot cannot be resumed safely. OpenAI's resumable RunState carries a schema version, while ADK's 2.0 release explicitly documents breaking changes; neither fact is a promise that every old snapshot will resume unchanged.
Conclusion
The best framework for a first agent is the one whose failure path your team can explain and operate. Start with a single task, two tools, one approval, and one forced restart. If the run remains authorized, reviewable, and recoverable, expand the workflow. If it does not, more agents will multiply the ambiguity.
