Human approval in production: LangGraph + FastAPI
In Part 1 the agent paused and asked in the console. In a real system the human is in Slack, in their inbox or on a separate screen — and may answer an hour later. Here we turn the agent's pause into a service: the graph stops with interrupt(), the pause is stored in a database, and FastAPI offers three endpoints — start, resume and state. Then we connect approval to n8n and gotoHuman.
interrupt_before but resumed with Command(resume=True); with a static pause that value goes nowhere, so "approved" never became true. AsyncPostgresSaver is opened with async with (not a plain call). Removed: CORS "*", returning the exception text to the client, an internal email address. Tracing now uses LANGSMITH_TRACING (the old name was LANGCHAIN_TRACING_V2). The n8n JSON used a non-existent chat node type and could not be imported — replaced with a table of the real nodes and the new in-Slack approval. n8n now has a built-in human review for AI Agent tools — the comparison with gotoHuman is updated. The gotoHuman YAML "template" is not a real format — reviews are built in a visual editor there. The example is the sick-leave case from Part 1 (paid by the social security institute, not the health fund), amounts in EUR.
01What you will learn
- Why human approval uses
interrupt()and not the static breakpointinterrupt_before. - How to let the human not only say "yes" or "no" but also fix the arguments before the action runs.
- How to expose the agent as an API with three endpoints:
/invoke,/resume,/state. - How to make the pause survive a restart — with PostgreSQL instead of memory.
- How to deliver the question to a human via n8n (Slack and other channels) or via gotoHuman.
- What to check before the agent meets real users.
02Before you start
- You have done Block 3 · Part 1 (tools,
StateGraph,interrupt()) — we reuse itscustom_tools.py. Part 2 (n8n) helps with the n8n section. - Python 3.11 or newer. We checked: on 3.10,
interrupt()in an async graph fails withCalled get_config outside of a runnable context. - Ollama 🔒 local with a model that supports tool calling (for example
llama3.1:8b). - For production: PostgreSQL 🔒 local (for example in Docker). Not required for the exercise.
python3.12 -m venv .venv && source .venv/bin/activate
pip install -U langgraph langchain langchain-ollama fastapi "uvicorn[standard]" \
langgraph-checkpoint-postgres "psycopg[binary,pool]"
# checked with: langgraph 1.2.12 · langgraph-checkpoint-postgres 3.1.2 · fastapi 0.142.2 · uvicorn 0.54.003Steps
-
Two kinds of pause — and why we use only one
LangGraph can stop a graph in two ways. Why does it matter? They look alike, but only one carries the human's decision back into the graph.
Pause How it works What for How it resumes interrupt()in a nodeStops anywhere in the code, conditionally. Shows the human whatever you pass it. Human approval — the recommended way. Command(resume=value)→ the value is returned byinterrupt()compile(interrupt_before=[…])Static breakpoint: always stops before the given node. Debugging, step by step. invoke(None, config)— no value⚠️The trap in the old version of this lessonThe graph paused withinterrupt_beforeand resumed withCommand(resume=True). With a static pause there is nointerrupt()to receive that value — it is simply lost. The "approved" field stayed empty and the action never ran. The official docs are clear: static interrupts are not recommended for human approval. -
The graph: a pause that allows edits
We take the graph from Part 1 and extend the
human_checknode. The human returns a decision with three fields:approved, optionallyedited_args(corrected arguments) andcomment. Why edits? The agent is often almost right — wrong recipient, rounded amount. It is faster for the human to fix it than to reject and let the agent try again.Python · hil_agent.pyfrom typing import Annotated, Literal, TypedDict from langchain_core.messages import AnyMessage, ToolMessage from langchain_ollama import ChatOllama from langgraph.graph import StateGraph, START, END from langgraph.graph.message import add_messages from langgraph.prebuilt import ToolNode from langgraph.types import interrupt, Command from custom_tools import TOOLS, ACTION_TOOLS # from Part 1 class AgentState(TypedDict): messages: Annotated[list[AnyMessage], add_messages] llm_with_tools = ChatOllama(model="llama3.1:8b", temperature=0.1).bind_tools(TOOLS) async def agent_node(state: AgentState) -> dict: return {"messages": [await llm_with_tools.ainvoke(state["messages"])]} def human_check(state: AgentState) -> Command[Literal["tools", "agent"]]: last = state["messages"][-1] actions = [{"id": tc["id"], "name": tc["name"], "args": tc["args"]} for tc in last.tool_calls if tc["name"] in ACTION_TOOLS] # PAUSE. Whatever you pass in Command(resume=...) becomes the value of decision. decision = interrupt({"type": "approval", "actions": actions}) if decision.get("approved"): edits = decision.get("edited_args") or {} if not edits: return Command(goto="tools") calls = [{**tc, "args": {**tc["args"], **edits}} if tc["name"] in ACTION_TOOLS else tc for tc in last.tool_calls] # same id → add_messages REPLACES the message instead of appending return Command(goto="tools", update={"messages": [last.model_copy(update={"tool_calls": calls})]}) # Reject: every tool call gets a reply, otherwise the next model call fails reason = decision.get("comment") or "Rejected by a human." rejected = [ToolMessage(content=f"{reason} Do not retry; explain to the user.", tool_call_id=tc["id"]) for tc in last.tool_calls] return Command(goto="agent", update={"messages": rejected}) def route(state: AgentState) -> Literal["tools", "human_check", "__end__"]: last = state["messages"][-1] if not getattr(last, "tool_calls", None): return END if any(tc["name"] in ACTION_TOOLS for tc in last.tool_calls): return "human_check" return "tools" def build_graph(checkpointer): """The graph receives its checkpointer from outside — in memory for practice, PostgreSQL for production.""" g = StateGraph(AgentState) g.add_node("agent", agent_node) g.add_node("tools", ToolNode(TOOLS)) g.add_node("human_check", human_check) g.add_edge(START, "agent") g.add_conditional_edges("agent", route, ["tools", "human_check", END]) g.add_edge("tools", "agent") return g.compile(checkpointer=checkpointer)⚠️Trap: the node starts overOn resume LangGraph runshuman_checkfrom its first line, not from theinterrupt()line. So beforeinterrupt()there is no writing, sending or paying — only reading. The action runs intools, after the decision. -
FastAPI: the agent as a service with three endpoints
/invokestarts a request and returns either an answer or "waiting for approval" with the details./resumetakes the decision./state/{thread_id}shows how far a thread got — handy for checks from n8n. Why separate endpoints? Hours may pass between the pause and the decision — an HTTP request cannot hang that long.Python · agent_api.pyimport logging, os, uuid from contextlib import asynccontextmanager from typing import Literal from fastapi import Depends, FastAPI, Header, HTTPException, Request from pydantic import BaseModel, Field from langchain_core.messages import HumanMessage from langgraph.checkpoint.memory import InMemorySaver from langgraph.types import Command from hil_agent import build_graph log = logging.getLogger("agent_api") DB_URI = os.getenv("DB_URI") # PostgreSQL connection — from the environment, never in code API_TOKEN = os.getenv("AGENT_API_TOKEN") # shared secret between n8n and the API @asynccontextmanager async def lifespan(app: FastAPI): if DB_URI: # production: pauses survive restarts from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver async with AsyncPostgresSaver.from_conn_string(DB_URI) as saver: await saver.setup() # creates the tables on first run app.state.graph = build_graph(saver) yield else: # practice: everything in memory app.state.graph = build_graph(InMemorySaver()) yield api = FastAPI(title="HIL Agent API", version="1.0", lifespan=lifespan) def auth(authorization: str = Header(default="")): if API_TOKEN and authorization != f"Bearer {API_TOKEN}": raise HTTPException(401, "unauthorized") class InvokeRequest(BaseModel): message: str = Field(min_length=1, max_length=4000) thread_id: str | None = Field(default=None, max_length=200) class ResumeRequest(BaseModel): thread_id: str = Field(max_length=200) approved: bool edited_args: dict | None = None comment: str | None = Field(default=None, max_length=1000) class AgentResponse(BaseModel): thread_id: str status: Literal["completed", "interrupted"] final_answer: str | None = None pending: dict | None = None # when interrupted: what awaits approval def cfg(thread_id: str) -> dict: return {"configurable": {"thread_id": thread_id}, "recursion_limit": 25} def to_response(thread_id: str, result: dict) -> AgentResponse: if "__interrupt__" in result: return AgentResponse(thread_id=thread_id, status="interrupted", pending=result["__interrupt__"][0].value) return AgentResponse(thread_id=thread_id, status="completed", final_answer=result["messages"][-1].content) @api.post("/invoke", response_model=AgentResponse, dependencies=[Depends(auth)]) async def invoke(req: InvokeRequest, request: Request): thread_id = req.thread_id or str(uuid.uuid4()) try: result = await request.app.state.graph.ainvoke( {"messages": [HumanMessage(content=req.message)]}, cfg(thread_id)) except Exception: log.exception("invoke failed") # details go to the log, not to the client raise HTTPException(500, "agent error") return to_response(thread_id, result) @api.post("/resume", response_model=AgentResponse, dependencies=[Depends(auth)]) async def resume(req: ResumeRequest, request: Request): graph = request.app.state.graph snap = await graph.aget_state(cfg(req.thread_id)) if not snap.interrupts: # no pending pause → nothing to resume raise HTTPException(409, "nothing to resume") decision = req.model_dump(include={"approved", "edited_args", "comment"}) try: result = await graph.ainvoke(Command(resume=decision), cfg(req.thread_id)) except Exception: log.exception("resume failed") raise HTTPException(500, "agent error") return to_response(req.thread_id, result) @api.get("/state/{thread_id}", dependencies=[Depends(auth)]) async def state(thread_id: str, request: Request): snap = await request.app.state.graph.aget_state(cfg(thread_id)) if not snap.values: raise HTTPException(404, "unknown thread") msgs = snap.values.get("messages", []) return {"thread_id": thread_id, "status": "interrupted" if snap.interrupts else "completed", "pending": snap.interrupts[0].value if snap.interrupts else None, "messages": len(msgs), "last_message": msgs[-1].content if msgs else None}bash · run and try by handexport AGENT_API_TOKEN="<long-random-secret>" # optional: export DB_URI="postgresql://<user>:<password>@<host>:5432/<db>" uvicorn agent_api:api --port 8001 # by default listens on this machine only # in another terminal curl -s localhost:8001/invoke -H "Authorization: Bearer $AGENT_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"message":"Tell HR that Ivan is on sick leave until Friday.","thread_id":"hr-1"}' # → {"status":"interrupted","pending":{"type":"approval","actions":[...]}, ...} curl -s localhost:8001/resume -H "Authorization: Bearer $AGENT_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"thread_id":"hr-1","approved":true,"edited_args":{"to":"hr@example.com"}}' # → {"status":"completed","final_answer":"..."}✅Whyasync withand lifespanAsyncPostgresSaver.from_conn_string(…)opens a connection and must close it — that is why it is a context manager. In FastAPI it belongs inlifespan: opened once at startup, closed at shutdown.setup()creates the tables; it is safe to call on every start.⛔Do not expose the API bare to the internetWithout a key, anyone can trigger an action or approve it in your place. The minimum: a secret in the header, HTTPS through a reverse proxy, noCORS "*", no exception text to the client (it can contain paths and connection details), rate limiting. -
n8n: the question reaches a human, the decision comes back
n8n is the "postman": it takes the user's question, calls
/invoke, on a pause sends the question to the approver and posts the decision to/resume. Why not Slack → API directly? Because n8n already has ready channels, retries and an execution history.# n8n node Setting 1 Chat Trigger The user's input. Use its sessionIdasthread_id.2 HTTP Request · POST /invokeJSON body: message={{ $json.chatInput }},thread_id={{ $json.sessionId }}. HeaderAuthorization: Bearer …— as a credential, not in plain text.3 IF {{ $json.status }}equalsinterrupted.4 Slack · Message → Send and Wait for Response Response Type: Approval. In the text — the tool and arguments from pending.actions. The approver clicks a button in Slack; the output containsapprovedand who responded. Instead of Slack you can use Chat, Teams, Telegram, Gmail and others.5 HTTP Request · POST /resumethread_idfrom step 2,approvedfrom step 4, optionallycomment.6 Chat · Send Message Returns final_answerto the user (Chat Trigger with Response Mode "Using Response Nodes"). The "false" branch of the IF goes straight here.🔄Approval in Slack — what you needn8n must be reachable from Slack over public HTTPS (it does not work withlocalhost). In the Slack app, turn on Interactivity with the URLhttps://<your-n8n>/webhook-waiting-slack, and put the Signing Secret into the Slack credential in n8n. Without it the buttons show but do not resume the flow.✅If the agent lives inside n8nWhen the agent is an n8n AI Agent node (not Python), you do not need FastAPI: the Tools panel has a Human review section. Pick a channel and connect under it the tools that need approval. The human can approve or deny (no editing); on deny the agent is told and explains. Describe in the system prompt which tools go through review. -
gotoHuman: a ready review screen
When there are several approvers who want a queue, value editing and history, a custom screen costs weeks. gotoHuman 🌐 global offers these as a service. Why consider it? You do not write a user interface — you only describe the fields.
Need n8n Human review (built-in) Slack approval + your API gotoHuman Approve / reject Yes Yes Yes Edit the arguments No Only if you build it ( edited_args)Yes, in the review fields Shared queue for the team No No Yes (Agent Inbox) Who approved Execution history Yes (Slack output) Yes Data leaves your machine Only via the chosen channel To Slack To gotoHuman steps · gotoHuman in n8n# 1. Install the node: n8n → nodes panel → search "gotoHuman" (verified community node) # Self-hosted n8n without the panel: Settings → Community Nodes → Install → @gotohuman/n8n-nodes-gotohuman # 2. Account at gotohuman.com → API key → in n8n: Credentials → gotoHuman API # 3. In the gotoHuman web editor: a new review type (fields + buttons) # 4. In the flow: gotoHuman node → Send and Wait for Response # The flow pauses and resumes by itself with the review result.Review field Type Editable Tool text no Why the agent wants this markdown no Average daily gross (EUR) number yes Working days of sick leave number yes Estimated amount (EUR) number yes Comment text yes, optional 💡This is a field plan, not a fileThe old lesson showed a YAML configuration. gotoHuman has no such format — a review type is assembled in the visual editor, and the result arrives in the node output or via webhook. See the documentation for the exact field types (link in Sources). For a Python agent without n8n, their docs include a LangGraph example.⚠️Personal dataSick leave is health data. Before sending such fields to an external service (Slack, gotoHuman) — a data processing agreement and a minimum of fields. Otherwise keep the review in your own n8n chat. -
Before production: what to check
An agent with access to actions is software with permissions. Every row skipped below is a real risk — a lost approval, an email sent twice, leaked data.
Priority Check high Every action with a real effect (send, write, update, delete) goes through interrupt().high A persistent checkpointer (PostgreSQL) — pauses survive restarts. high A step ceiling: recursion_limitin the config (LangGraph has nomax_iterations).high Tools return a structured error rather than raising to the agent. high API with a key, HTTPS, no CORS "*", no error text to the client.medium Tracing: LANGSMITH_TRACING=true+LANGSMITH_API_KEY🌐 or local logs.medium Rate limiting (slowapi or at the reverse proxy). medium A unit test per tool: normal input, error, edge case. medium Tool descriptions tested with hard questions — the agent picks the right one. medium A system prompt with clear limits: who the agent is, what it does NOT do, when it asks a human. medium Token monitoring — agents spend a lot without it showing. low Input and output guardrails (NeMo Guardrails or LLM Guard) for health and legal topics. low An isolated sandbox for tools that execute code. low Compare models for Bulgarian (e.g. Llama 3.1 vs BgGPT) on your own real questions. 🎯Which tool for whatLangGraph — the agent's logic and the pauses in critical processes (contracts, health documents). CrewAI — a quick prototype of an agent team. n8n — orchestration, notifications, approvals. AgentExecutor — only in legacy projects (now inlangchain-classic). Exercise: take a CrewAI team from Part 2 → move the logic into a LangGraph graph → put FastAPI in front → connect it to n8n.
04Check
Checklist
- Python is 3.11+; imports of
interrupt,CommandandAsyncPostgresSaversucceed. /invokewith a request for an action returns"status": "interrupted"with the details inpending./resumewithapproved: trueruns the action; withedited_args— with the corrected values./resumewithapproved: false— the agent explains the rejection, nothing is sent.- A second
/resumeon the same thread → 409; unknown thread on/state→ 404; no key → 401. - With
DB_URI: stop the API during a pause, start it again —/resumecontinues. - The n8n approval shows the tool and arguments, and the decision reaches
/resume.
Quiz
1. The graph is compiled with interrupt_before=["human_check"] and resumed with Command(resume=True). What happens?
2. Why is there no email sending in human_check before interrupt()?
3. The API restarts while waiting for approval. When will /resume succeed?
4. Your agent is an AI Agent node in n8n and you want approval only for "send email". What is simplest?
05What's next
06Sources
- LangGraph: interrupts —
interrupt(),Command(resume), the re-run rules, static interrupts for debugging only. - LangGraph: persistence — checkpointers, threads, the
thread_idlength limit. - LangGraph: memory with PostgreSQL —
AsyncPostgresSaverwithasync withandsetup(). - langgraph-checkpoint-postgres on PyPI and fastapi on PyPI — current versions.
- FastAPI: lifespan — resources opened at startup and closed at shutdown.
- LangSmith: tracing LangGraph 🌐 global —
LANGSMITH_TRACING,LANGSMITH_API_KEY. - n8n: Human-in-the-loop for tools — the built-in review in AI Agent.
- n8n: approvals in Slack — Send and Wait for Response, Interactivity, Signing Secret.
- n8n: Chat Trigger and Chat — input and reply in the chat.
- gotoHuman: n8n, review types and LangGraph example 🌐 global.
- slowapi — rate limiting for FastAPI.
- NeMo Guardrails and LLM Guard — input and output guardrails.