The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-03c · Human approval in production: LangGraph + FastAPI
series
Blocks 0–10 · Block 3 — Agents and orchestration · Appendix (after parts 1 and 2)
level
Advanced
duration
5–8 h
prerequisites
01-03a (tools, StateGraph, interrupt) · 01-03b (n8n as orchestrator); Python ≥ 3.11 for async interrupt()
trust_label
VERIFIED 2026-10-01 (package versions on PyPI/npm, API names in official docs, FastAPI + graph executed with a fake chat model via TestClient: invoke → interrupted → resume approve/edit/reject → completed, 401/409/404 paths; all source links) · UPDATED 2026-10-01 · NOT end-to-end tested with a real LLM, PostgreSQL, n8n or gotoHuman
versions
langgraph 1.2.x · langgraph-checkpoint-postgres 3.1.x · langchain 1.4.x · fastapi 0.142.x · uvicorn 0.54.x · n8n 2.41.x · @gotohuman/n8n-nodes-gotohuman 0.4.x
language
human view: en (this edition) · bulgarian edition: /academy/blokove/moduli/01-03c_Блок_3_Приложение_HIL_FastAPI.html
next
01-04a_Блок_4_Част_1_Advanced_RAG.html · Advanced RAG
PURPOSE

Turn the approval pause of a LangGraph agent into a service. The graph pauses with interrupt() before any action tool, the pause is persisted by a checkpointer, and a FastAPI app exposes /invoke (start), /resume (approve, edit arguments or reject) and /state (inspect). n8n or gotoHuman delivers the question to a human and posts the decision back. Finish with a production hardening checklist.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-04a_Блок_4_Част_1_Advanced_RAG.html · Advanced RAG · previous: 01-03b_Блок_3_Част_2_CrewAI_n8n.html · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
human-in-the-looplanggraphinterruptfastapipostgrescheckpointern8napprovalsgotohumanproduction
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Human approval in production: LangGraph + FastAPI

In Part 1 the agent paused and asked in the console. In a real system the human is in Slack, in their inbox or on a separate screen — and may answer an hour later. Here we turn the agent's pause into a service: the graph stops with interrupt(), the pause is stored in a database, and FastAPI offers three endpoints — start, resume and state. Then we connect approval to n8n and gotoHuman.

⏱ 5–8 h Advanced Block 3 · Appendix approval · API · production
LangGraph · FastAPI🔒 local PostgreSQL (memory of pauses)🔒 local Ollama (the model)🔒 local n8n (self-hosted)🔒 local Slack · gotoHuman · LangSmith🌐 global
🔄
UPDATED · 01.10.2026 — what changed
The code is rewritten for LangGraph 1.2 and FastAPI 0.142 and was executed with a fake model: start → pause → approve / edit / reject → done. Bugs fixed from the old version: the pause used interrupt_before but resumed with Command(resume=True); with a static pause that value goes nowhere, so "approved" never became true. AsyncPostgresSaver is opened with async with (not a plain call). Removed: CORS "*", returning the exception text to the client, an internal email address. Tracing now uses LANGSMITH_TRACING (the old name was LANGCHAIN_TRACING_V2). The n8n JSON used a non-existent chat node type and could not be imported — replaced with a table of the real nodes and the new in-Slack approval. n8n now has a built-in human review for AI Agent tools — the comparison with gotoHuman is updated. The gotoHuman YAML "template" is not a real format — reviews are built in a visual editor there. The example is the sick-leave case from Part 1 (paid by the social security institute, not the health fund), amounts in EUR.

01What you will learn

02Before you start

bash · install (versions as of 01.10.2026)
python3.12 -m venv .venv && source .venv/bin/activate
pip install -U langgraph langchain langchain-ollama fastapi "uvicorn[standard]" \
               langgraph-checkpoint-postgres "psycopg[binary,pool]"
# checked with: langgraph 1.2.12 · langgraph-checkpoint-postgres 3.1.2 · fastapi 0.142.2 · uvicorn 0.54.0

03Steps

  1. Two kinds of pause — and why we use only one

    LangGraph can stop a graph in two ways. Why does it matter? They look alike, but only one carries the human's decision back into the graph.

    PauseHow it worksWhat forHow it resumes
    interrupt() in a nodeStops anywhere in the code, conditionally. Shows the human whatever you pass it.Human approval — the recommended way.Command(resume=value) → the value is returned by interrupt()
    compile(interrupt_before=[…])Static breakpoint: always stops before the given node.Debugging, step by step.invoke(None, config) — no value
    ⚠️
    The trap in the old version of this lesson
    The graph paused with interrupt_before and resumed with Command(resume=True). With a static pause there is no interrupt() to receive that value — it is simply lost. The "approved" field stayed empty and the action never ran. The official docs are clear: static interrupts are not recommended for human approval.
  2. The graph: a pause that allows edits

    We take the graph from Part 1 and extend the human_check node. The human returns a decision with three fields: approved, optionally edited_args (corrected arguments) and comment. Why edits? The agent is often almost right — wrong recipient, rounded amount. It is faster for the human to fix it than to reject and let the agent try again.

    Python · hil_agent.py
    from typing import Annotated, Literal, TypedDict
    from langchain_core.messages import AnyMessage, ToolMessage
    from langchain_ollama import ChatOllama
    from langgraph.graph import StateGraph, START, END
    from langgraph.graph.message import add_messages
    from langgraph.prebuilt import ToolNode
    from langgraph.types import interrupt, Command
    from custom_tools import TOOLS, ACTION_TOOLS          # from Part 1
    
    class AgentState(TypedDict):
        messages: Annotated[list[AnyMessage], add_messages]
    
    llm_with_tools = ChatOllama(model="llama3.1:8b", temperature=0.1).bind_tools(TOOLS)
    
    async def agent_node(state: AgentState) -> dict:
        return {"messages": [await llm_with_tools.ainvoke(state["messages"])]}
    
    def human_check(state: AgentState) -> Command[Literal["tools", "agent"]]:
        last = state["messages"][-1]
        actions = [{"id": tc["id"], "name": tc["name"], "args": tc["args"]}
                   for tc in last.tool_calls if tc["name"] in ACTION_TOOLS]
        # PAUSE. Whatever you pass in Command(resume=...) becomes the value of decision.
        decision = interrupt({"type": "approval", "actions": actions})
        if decision.get("approved"):
            edits = decision.get("edited_args") or {}
            if not edits:
                return Command(goto="tools")
            calls = [{**tc, "args": {**tc["args"], **edits}} if tc["name"] in ACTION_TOOLS else tc
                     for tc in last.tool_calls]
            # same id → add_messages REPLACES the message instead of appending
            return Command(goto="tools",
                           update={"messages": [last.model_copy(update={"tool_calls": calls})]})
        # Reject: every tool call gets a reply, otherwise the next model call fails
        reason = decision.get("comment") or "Rejected by a human."
        rejected = [ToolMessage(content=f"{reason} Do not retry; explain to the user.",
                                tool_call_id=tc["id"]) for tc in last.tool_calls]
        return Command(goto="agent", update={"messages": rejected})
    
    def route(state: AgentState) -> Literal["tools", "human_check", "__end__"]:
        last = state["messages"][-1]
        if not getattr(last, "tool_calls", None):
            return END
        if any(tc["name"] in ACTION_TOOLS for tc in last.tool_calls):
            return "human_check"
        return "tools"
    
    def build_graph(checkpointer):
        """The graph receives its checkpointer from outside — in memory for practice, PostgreSQL for production."""
        g = StateGraph(AgentState)
        g.add_node("agent", agent_node)
        g.add_node("tools", ToolNode(TOOLS))
        g.add_node("human_check", human_check)
        g.add_edge(START, "agent")
        g.add_conditional_edges("agent", route, ["tools", "human_check", END])
        g.add_edge("tools", "agent")
        return g.compile(checkpointer=checkpointer)
    ⚠️
    Trap: the node starts over
    On resume LangGraph runs human_check from its first line, not from the interrupt() line. So before interrupt() there is no writing, sending or paying — only reading. The action runs in tools, after the decision.
  3. FastAPI: the agent as a service with three endpoints

    /invoke starts a request and returns either an answer or "waiting for approval" with the details. /resume takes the decision. /state/{thread_id} shows how far a thread got — handy for checks from n8n. Why separate endpoints? Hours may pass between the pause and the decision — an HTTP request cannot hang that long.

    Python · agent_api.py
    import logging, os, uuid
    from contextlib import asynccontextmanager
    from typing import Literal
    from fastapi import Depends, FastAPI, Header, HTTPException, Request
    from pydantic import BaseModel, Field
    from langchain_core.messages import HumanMessage
    from langgraph.checkpoint.memory import InMemorySaver
    from langgraph.types import Command
    from hil_agent import build_graph
    
    log = logging.getLogger("agent_api")
    DB_URI = os.getenv("DB_URI")                  # PostgreSQL connection — from the environment, never in code
    API_TOKEN = os.getenv("AGENT_API_TOKEN")      # shared secret between n8n and the API
    
    @asynccontextmanager
    async def lifespan(app: FastAPI):
        if DB_URI:                                # production: pauses survive restarts
            from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
            async with AsyncPostgresSaver.from_conn_string(DB_URI) as saver:
                await saver.setup()               # creates the tables on first run
                app.state.graph = build_graph(saver)
                yield
        else:                                     # practice: everything in memory
            app.state.graph = build_graph(InMemorySaver())
            yield
    
    api = FastAPI(title="HIL Agent API", version="1.0", lifespan=lifespan)
    
    def auth(authorization: str = Header(default="")):
        if API_TOKEN and authorization != f"Bearer {API_TOKEN}":
            raise HTTPException(401, "unauthorized")
    
    class InvokeRequest(BaseModel):
        message: str = Field(min_length=1, max_length=4000)
        thread_id: str | None = Field(default=None, max_length=200)
    
    class ResumeRequest(BaseModel):
        thread_id: str = Field(max_length=200)
        approved: bool
        edited_args: dict | None = None
        comment: str | None = Field(default=None, max_length=1000)
    
    class AgentResponse(BaseModel):
        thread_id: str
        status: Literal["completed", "interrupted"]
        final_answer: str | None = None
        pending: dict | None = None               # when interrupted: what awaits approval
    
    def cfg(thread_id: str) -> dict:
        return {"configurable": {"thread_id": thread_id}, "recursion_limit": 25}
    
    def to_response(thread_id: str, result: dict) -> AgentResponse:
        if "__interrupt__" in result:
            return AgentResponse(thread_id=thread_id, status="interrupted",
                                 pending=result["__interrupt__"][0].value)
        return AgentResponse(thread_id=thread_id, status="completed",
                             final_answer=result["messages"][-1].content)
    
    @api.post("/invoke", response_model=AgentResponse, dependencies=[Depends(auth)])
    async def invoke(req: InvokeRequest, request: Request):
        thread_id = req.thread_id or str(uuid.uuid4())
        try:
            result = await request.app.state.graph.ainvoke(
                {"messages": [HumanMessage(content=req.message)]}, cfg(thread_id))
        except Exception:
            log.exception("invoke failed")        # details go to the log, not to the client
            raise HTTPException(500, "agent error")
        return to_response(thread_id, result)
    
    @api.post("/resume", response_model=AgentResponse, dependencies=[Depends(auth)])
    async def resume(req: ResumeRequest, request: Request):
        graph = request.app.state.graph
        snap = await graph.aget_state(cfg(req.thread_id))
        if not snap.interrupts:                   # no pending pause → nothing to resume
            raise HTTPException(409, "nothing to resume")
        decision = req.model_dump(include={"approved", "edited_args", "comment"})
        try:
            result = await graph.ainvoke(Command(resume=decision), cfg(req.thread_id))
        except Exception:
            log.exception("resume failed")
            raise HTTPException(500, "agent error")
        return to_response(req.thread_id, result)
    
    @api.get("/state/{thread_id}", dependencies=[Depends(auth)])
    async def state(thread_id: str, request: Request):
        snap = await request.app.state.graph.aget_state(cfg(thread_id))
        if not snap.values:
            raise HTTPException(404, "unknown thread")
        msgs = snap.values.get("messages", [])
        return {"thread_id": thread_id,
                "status": "interrupted" if snap.interrupts else "completed",
                "pending": snap.interrupts[0].value if snap.interrupts else None,
                "messages": len(msgs),
                "last_message": msgs[-1].content if msgs else None}
    bash · run and try by hand
    export AGENT_API_TOKEN="<long-random-secret>"
    # optional: export DB_URI="postgresql://<user>:<password>@<host>:5432/<db>"
    uvicorn agent_api:api --port 8001          # by default listens on this machine only
    
    # in another terminal
    curl -s localhost:8001/invoke -H "Authorization: Bearer $AGENT_API_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"message":"Tell HR that Ivan is on sick leave until Friday.","thread_id":"hr-1"}'
    # → {"status":"interrupted","pending":{"type":"approval","actions":[...]}, ...}
    
    curl -s localhost:8001/resume -H "Authorization: Bearer $AGENT_API_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"thread_id":"hr-1","approved":true,"edited_args":{"to":"hr@example.com"}}'
    # → {"status":"completed","final_answer":"..."}
    ✅
    Why async with and lifespan
    AsyncPostgresSaver.from_conn_string(…) opens a connection and must close it — that is why it is a context manager. In FastAPI it belongs in lifespan: opened once at startup, closed at shutdown. setup() creates the tables; it is safe to call on every start.
    ⛔
    Do not expose the API bare to the internet
    Without a key, anyone can trigger an action or approve it in your place. The minimum: a secret in the header, HTTPS through a reverse proxy, no CORS "*", no exception text to the client (it can contain paths and connection details), rate limiting.
  4. n8n: the question reaches a human, the decision comes back

    n8n is the "postman": it takes the user's question, calls /invoke, on a pause sends the question to the approver and posts the decision to /resume. Why not Slack → API directly? Because n8n already has ready channels, retries and an execution history.

    #n8n nodeSetting
    1Chat TriggerThe user's input. Use its sessionId as thread_id.
    2HTTP Request · POST /invokeJSON body: message = {{ $json.chatInput }}, thread_id = {{ $json.sessionId }}. Header Authorization: Bearer … — as a credential, not in plain text.
    3IF{{ $json.status }} equals interrupted.
    4Slack · Message → Send and Wait for ResponseResponse Type: Approval. In the text — the tool and arguments from pending.actions. The approver clicks a button in Slack; the output contains approved and who responded. Instead of Slack you can use Chat, Teams, Telegram, Gmail and others.
    5HTTP Request · POST /resumethread_id from step 2, approved from step 4, optionally comment.
    6Chat · Send MessageReturns final_answer to the user (Chat Trigger with Response Mode "Using Response Nodes"). The "false" branch of the IF goes straight here.
    🔄
    Approval in Slack — what you need
    n8n must be reachable from Slack over public HTTPS (it does not work with localhost). In the Slack app, turn on Interactivity with the URL https://<your-n8n>/webhook-waiting-slack, and put the Signing Secret into the Slack credential in n8n. Without it the buttons show but do not resume the flow.
    ✅
    If the agent lives inside n8n
    When the agent is an n8n AI Agent node (not Python), you do not need FastAPI: the Tools panel has a Human review section. Pick a channel and connect under it the tools that need approval. The human can approve or deny (no editing); on deny the agent is told and explains. Describe in the system prompt which tools go through review.
  5. gotoHuman: a ready review screen

    When there are several approvers who want a queue, value editing and history, a custom screen costs weeks. gotoHuman 🌐 global offers these as a service. Why consider it? You do not write a user interface — you only describe the fields.

    Needn8n Human review (built-in)Slack approval + your APIgotoHuman
    Approve / rejectYesYesYes
    Edit the argumentsNoOnly if you build it (edited_args)Yes, in the review fields
    Shared queue for the teamNoNoYes (Agent Inbox)
    Who approvedExecution historyYes (Slack output)Yes
    Data leaves your machineOnly via the chosen channelTo SlackTo gotoHuman
    steps · gotoHuman in n8n
    # 1. Install the node: n8n → nodes panel → search "gotoHuman" (verified community node)
    #    Self-hosted n8n without the panel: Settings → Community Nodes → Install →
    @gotohuman/n8n-nodes-gotohuman
    # 2. Account at gotohuman.com → API key → in n8n: Credentials → gotoHuman API
    # 3. In the gotoHuman web editor: a new review type (fields + buttons)
    # 4. In the flow: gotoHuman node → Send and Wait for Response
    #    The flow pauses and resumes by itself with the review result.
    Review fieldTypeEditable
    Tooltextno
    Why the agent wants thismarkdownno
    Average daily gross (EUR)numberyes
    Working days of sick leavenumberyes
    Estimated amount (EUR)numberyes
    Commenttextyes, optional
    💡
    This is a field plan, not a file
    The old lesson showed a YAML configuration. gotoHuman has no such format — a review type is assembled in the visual editor, and the result arrives in the node output or via webhook. See the documentation for the exact field types (link in Sources). For a Python agent without n8n, their docs include a LangGraph example.
    ⚠️
    Personal data
    Sick leave is health data. Before sending such fields to an external service (Slack, gotoHuman) — a data processing agreement and a minimum of fields. Otherwise keep the review in your own n8n chat.
  6. Before production: what to check

    An agent with access to actions is software with permissions. Every row skipped below is a real risk — a lost approval, an email sent twice, leaked data.

    PriorityCheck
    highEvery action with a real effect (send, write, update, delete) goes through interrupt().
    highA persistent checkpointer (PostgreSQL) — pauses survive restarts.
    highA step ceiling: recursion_limit in the config (LangGraph has no max_iterations).
    highTools return a structured error rather than raising to the agent.
    highAPI with a key, HTTPS, no CORS "*", no error text to the client.
    mediumTracing: LANGSMITH_TRACING=true + LANGSMITH_API_KEY 🌐 or local logs.
    mediumRate limiting (slowapi or at the reverse proxy).
    mediumA unit test per tool: normal input, error, edge case.
    mediumTool descriptions tested with hard questions — the agent picks the right one.
    mediumA system prompt with clear limits: who the agent is, what it does NOT do, when it asks a human.
    mediumToken monitoring — agents spend a lot without it showing.
    lowInput and output guardrails (NeMo Guardrails or LLM Guard) for health and legal topics.
    lowAn isolated sandbox for tools that execute code.
    lowCompare models for Bulgarian (e.g. Llama 3.1 vs BgGPT) on your own real questions.
    🎯
    Which tool for what
    LangGraph — the agent's logic and the pauses in critical processes (contracts, health documents). CrewAI — a quick prototype of an agent team. n8n — orchestration, notifications, approvals. AgentExecutor — only in legacy projects (now in langchain-classic). Exercise: take a CrewAI team from Part 2 → move the logic into a LangGraph graph → put FastAPI in front → connect it to n8n.

04Check

Checklist

Quiz

1. The graph is compiled with interrupt_before=["human_check"] and resumed with Command(resume=True). What happens?

2. Why is there no email sending in human_check before interrupt()?

3. The API restarts while waiting for approval. When will /resume succeed?

4. Your agent is an AI Agent node in n8n and you want approval only for "send email". What is simplest?

05What's next

06Sources

  1. LangGraph: interrupts — interrupt(), Command(resume), the re-run rules, static interrupts for debugging only.
  2. LangGraph: persistence — checkpointers, threads, the thread_id length limit.
  3. LangGraph: memory with PostgreSQL — AsyncPostgresSaver with async with and setup().
  4. langgraph-checkpoint-postgres on PyPI and fastapi on PyPI — current versions.
  5. FastAPI: lifespan — resources opened at startup and closed at shutdown.
  6. LangSmith: tracing LangGraph 🌐 global — LANGSMITH_TRACING, LANGSMITH_API_KEY.
  7. n8n: Human-in-the-loop for tools — the built-in review in AI Agent.
  8. n8n: approvals in Slack — Send and Wait for Response, Interactivity, Signing Secret.
  9. n8n: Chat Trigger and Chat — input and reply in the chat.
  10. gotoHuman: n8n, review types and LangGraph example 🌐 global.
  11. slowapi — rate limiting for FastAPI.
  12. NeMo Guardrails and LLM Guard — input and output guardrails.