Agents

Which should you choose: CrewAI, AutoGen or LangGraph?

CrewAI models role-based crews and event-driven flows, and is quickest for team-shaped tasks. AutoGen pioneered conversational multi-agent orchestration; its successor, Microsoft Agent Framework, continues that line with graph workflows. LangGraph is the most explicit: a stateful graph with durable execution, checkpointing and human-in-the-loop. All three are open source and bring-your-own-model, so the model layer can come from one OpenAI-compatible API.

Key facts

CrewAIPython framework: role-based crews plus event-driven flows
AutoGenConversational multi-agent framework; succeeded by Microsoft Agent Framework
Microsoft Agent FrameworkMerges AutoGen and Semantic Kernel ideas with graph-based workflows
LangGraphLow-level graph runtime with durable execution and checkpointers
State and resumeLangGraph checkpoints threads; Agent Framework adds checkpointed workflows; CrewAI flows keep explicit state
Model layerAll three are bring-your-own; Plugsky serves 30+ models on one key
Human in the loopNative interrupts in LangGraph; the others via custom gates
Product statusFramework choice is independent of the Plugsky API, which is live

TL;DR

  • Pick by programming model, not by model support, which is portable.
  • CrewAI for role-based teams, fastest path to a working crew.
  • AutoGen's lineage continues in Microsoft Agent Framework as it evolves.
  • LangGraph when you need explicit state, resumability and interrupts.
  • Keep the model layer on one OpenAI-compatible API so frameworks stay swappable.

How it works, step by step

  1. Write down the workflow shape: team of roles, conversation, or explicit state graph.
  2. List the hard requirements: resumability, human approval, parallelism, long runs.
  3. Prototype the same small task in two frameworks with the same model.
  4. Compare how each one handles state, retries and failure of a single step.
  5. Run both prototypes against one frozen evaluation set and measure cost per task.
  6. Choose the framework with the least code you would have to rewrite later.
  7. Point the winner at a single OpenAI-compatible endpoint to keep the model layer portable.
1Write down theworkflow shape:team of roles,2List the hardrequirements:resumability, human3Prototype the samesmall task in twoframeworks with the4Compare how eachone handles state,retries and failure5Run both prototypesagainst one frozenevaluation set and6Choose theframework with theleast code you

Try it yourself

Open the multi-agent workflow tool →

How each framework models the work

CrewAI starts from people: each agent has a role, a goal and a backstory, and the crew executes tasks sequentially or hierarchically. Flows add event-driven control when you need determinism. It is the shortest path from an idea to a working multi-agent prototype.

AutoGen started from conversations: agents exchange messages under termination conditions, with an event-driven core and higher-level team abstractions. That lineage now continues in Microsoft Agent Framework, which merges AutoGen's agent model with Semantic Kernel's enterprise features and adds typed, graph-based workflows with checkpointing.

LangGraph starts from state: you define nodes and edges over a state object, compile a graph, and attach a checkpointer. Every super-step is persisted against a thread, which enables resumption, time travel and interrupts. It is the lowest-level of the three and the most explicit.

State, failure and human-in-the-loop

These differences matter in production. LangGraph's checkpointers persist state at each step, so a failed node resumes from the last successful one rather than restarting the run; durability modes let you trade write latency against crash safety. Interrupts pause a graph for a human to inspect or edit state, then resume the same thread.

  • Long-running agents: LangGraph is built for resumable, stateful execution.
  • Conversation-centric tasks: AutoGen-style and Agent Framework patterns fit well.
  • Role-based pipelines: CrewAI expresses them with the least ceremony.
  • Human approval: native as an interrupt in LangGraph; a custom gate elsewhere.

In every case, evaluation and observability are your responsibility, and each ecosystem has its own tracing options.

The model layer is a separate decision

None of these frameworks is tied to one vendor's models; all three accept OpenAI-format endpoints alongside other providers. That means the framework decision should be about orchestration, and the model decision about quality, price and deployment. Keeping them separate is what protects you when either side changes.

Plugsky serves that model layer: 30+ models on one OpenAI-compatible key with live function calling, streaming, JSON mode, embeddings and RAG, plus scoped keys, RBAC, SSO/SCIM and audit logs, and deployment from shared cloud to VPC, on-prem and air-gapped. Framework-side features such as code execution sandboxes and tracing dashboards are not included, and assistants, files, batch and fine-tuning endpoints are coming soon. Plans and the free plugsky-micro and plugsky-lite models are on the live pricing page.

Honest comparison

DimensionCrewAIAutoGen / Agent FrameworkLangGraph
Core modelRoles, crews and flowsConversations and graph workflowsState graph of nodes and edges
Learning curveGentleModerateSteep but explicit
State and resumeExplicit state in flowsSession state, checkpointed workflowsCheckpointers per thread
Human in the loopCustom gateCustom gateNative interrupts
Best fitTeam-shaped tasksResearch and evolving patternsLong-running production agents

Frequently asked questions

Which framework is fastest to learn?

CrewAI, because roles and tasks map to how teams already describe work. LangGraph takes longer but gives finer control over state and resumption.

Which handles long-running agents best?

LangGraph, thanks to checkpointers, threads and durability modes that persist state at every step and resume after failure.

Is AutoGen still maintained?

AutoGen continues to receive critical fixes while Microsoft Agent Framework, built by the same teams, becomes the successor with graph-based workflows and migration guides.

Can I use these frameworks with non-OpenAI models?

Yes. All three are bring-your-own-model and accept OpenAI-compatible endpoints, so one API can serve many frameworks.

Do I need a framework at all?

No. A single tool-calling loop with your own state is often enough, and easier to debug. Reach for a framework when you need delegation, resumability or explicit graphs.

How do I avoid framework lock-in?

Keep tools and state provider-agnostic, keep the model layer behind an OpenAI-compatible interface, and keep evaluation data in a format the next framework can use.

Does Plugsky replace any of them?

No. Plugsky is the model and platform layer. It replaces the provider integration, not the orchestration framework.