代码仓库ChainReaction
给 agent 接工具,真正花时间的部分往往不是写函数体。函数体十行就写完了,剩下的时间花在接口对齐上:这个数据源要什么参数,那个数据源返回的结构怎么塞给模型,调用失败时错误信息用什么格式回传。每接一个系统,这套活儿重来一遍。
MCP(Model Context Protocol)把这层约定协议化了。服务端自己声明它有哪些工具、每个工具叫什么、参数 schema 长什么样;客户端连上去把这些声明读出来,转成模型能看懂的格式。适配代码从「每个数据源写一遍」变成「一份通用实现」。
LangChain 从 1.4 开始把 MCP 客户端收进了主包,命名空间是 langchain.mcp,核心类叫 MCPAdapter。它在 FastMCP 之上做转换:发现服务端的工具,变成普通的 LangChain 工具,直接丢给 create_agent。传输、协议握手、认证这些往下的事由 FastMCP 负责。
这篇按能跑的顺序讲:调用链长什么样,stdio 和 HTTP 两种连法差在哪,多个 server 挤在一起时工具名怎么处理,连接什么时候开什么时候关,认证怎么挂,服务端反过来要客户端能力时怎么办,手边一个 MCP server 都没有时怎么本地自测。文中的输出都是在本机跑出来的,脚本放在 MCP/ 。
MCP 标准化了哪一层 一个 MCP server 对外暴露的东西很薄:工具名、描述、输入 JSON Schema,外加一组可选的注解(比如 readOnlyHint、destructiveHint)。客户端用 tools/list 把这份目录读回来,用 tools/call 执行,返回值里带内容块、可选的 structuredContent,以及一个 isError 标志。
就这么几样东西,但它把「怎么发现工具」和「怎么调用工具」固定下来了。以前你为 GitHub 写一个 search_repos 工具,为内部知识库写一个 search_docs 工具,两份代码里的参数校验、错误处理、返回值包装各写各的。现在只要对方提供了 MCP server,你换一个 target 就完事,agent 那边的代码一行不动。
代价是引入了一层间接。工具的 schema 由服务端定义,你不再能随手改描述文案;服务端挂了,你的 agent 就少一块能力。这笔账划不划算,取决于你要接多少个外部系统。接三个以内,自己写工具更快;接十个,MCP 省下来的维护成本就很明显了。
调用链与接入点
sequenceDiagram
participant U as 用户
participant A as LangChain Agent
participant T as MCPAdapter 适配出的 Tool
participant C as FastMCP Client
participant S as MCP Server
U->>A: ainvoke(messages)
A->>A: 模型决定调用哪个工具
A->>T: 按普通 LangChain Tool 调用
T->>C: 进入 client 上下文并 tools/call
C->>S: stdio 子进程 或 Streamable HTTP
S-->>C: content + structuredContent + isError
C-->>T: CallToolResult
T-->>A: ToolMessage(status, content, artifact)
A->>U: 生成最终回复
从 agent 的角度看,MCP 工具和 @tool 装饰出来的工具没有区别。多出来的两层是 MCPAdapter 做的转换和 FastMCP 做的传输。
装依赖时注意,MCP 是主包的额外依赖:
1 pip install "langchain[mcp]"
这个命名空间还处于 beta,导入时会抛一次 LangChainBetaWarning。API 可能变,生产环境锁版本。
最小用法是三步:开 adapter、列工具、建 agent。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 import asyncioimport osfrom langchain.agents import create_agentfrom langchain.mcp import MCPAdapterfrom langchain_openai import ChatOpenAImodel = ChatOpenAI( api_key=os.getenv("DEEPSEEK_API_KEY" ), base_url="https://api.deepseek.com/v1" , model="deepseek-chat" , temperature=0.1 , max_tokens=1000 , ) async def main (): async with MCPAdapter("https://example.com/mcp" ) as adapter: tools = await adapter.list_tools() agent = create_agent(model, tools) result = await agent.ainvoke( {"messages" : [{"role" : "user" , "content" : "查一下 Oslo 的天气" }]} ) return result asyncio.run(main())
没有现成 server 也能验证这条链。下面这个 server 只有三十行,用 FastMCP 写,跑 stdio。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 from fastmcp import FastMCPmcp = FastMCP("weather-lab" , instructions="演示用天气 / 算术 MCP server" ) _FORECAST = { "Oslo" : {"temp_c" : 3 , "summary" : "cloudy" }, "Shanghai" : {"temp_c" : 19 , "summary" : "light rain" }, } @mcp.tool(annotations={"readOnlyHint" : True , "title" : "查天气" } ) def get_forecast (city: str ) -> dict : """Return a short weather forecast for a city.""" data = _FORECAST.get(city, {"temp_c" : 20 , "summary" : "unknown city, fallback" }) return {"city" : city, **data} @mcp.tool(annotations={"readOnlyHint" : True } ) def add (a: float , b: float ) -> float : """Add two numbers.""" return a + b @mcp.tool(annotations={"destructiveHint" : True , "readOnlyHint" : False } ) def divide (a: float , b: float ) -> float : """Divide a by b. Fails when b is zero.""" if b == 0 : raise ValueError("division by zero" ) return a / b if __name__ == "__main__" : mcp.run(transport="stdio" , show_banner=False )
连上去列一遍工具,看看适配结果长什么样。我用 Path 指脚本,走 stdio:
1 2 3 4 SERVER = Path(__file__).with_name("mcp_weather_server.py" ) async with MCPAdapter(SERVER) as adapter: tools = await adapter.list_tools()
真实输出(省略了 add 那条,结构和 get_forecast 一样):
1 2 3 4 5 6 - get_forecast: Return a short weather forecast for a city. args_schema: {"type": "object", "additionalProperties": false, "properties": {"city": {"type": "string"}}, "required": ["city"]} metadata: {"mcp": {"tool": {"annotations": {"title": "查天气", "read_only_hint": true}, "_meta": {"fastmcp": {"tags": []}}}, "server": {"name": "weather-lab", "version": "4.0.10"}}} - divide: Divide a by b. Fails when b is zero. args_schema: {"type": "object", "additionalProperties": false, "properties": {"a": {"type": "number"}, "b": {"type": "number"}}, "required": ["a", "b"]} metadata: {"mcp": {"tool": {"annotations": {"read_only_hint": false, "destructive_hint": true}, ...}, "server": {"name": "weather-lab", "version": "4.0.10"}}}
服务端写的注解原样落到了 LangChain 工具的 metadata["mcp"] 下面,readOnlyHint 变成 snake_case 的 read_only_hint,服务端身份在 mcp.server 里。这些字段全是可选的,读的时候要带默认值,不能假设一定存在。
不经过模型直接调工具,能看清返回值被拆成了什么。传一个完整的 ToolCall 进去:
1 2 3 msg = await forecast.ainvoke( {"type" : "tool_call" , "name" : "get_forecast" , "args" : {"city" : "Oslo" }, "id" : "call_1" } )
1 2 3 type=ToolMessage status=success content={"city":"Oslo","temp_c":3,"summary":"cloudy"} artifact={'structured_content': {'city': 'Oslo', 'temp_c': 3, 'summary': 'cloudy'}}
工具返回 dict,模型看到的是压成一行 JSON 的文本,结构化的那份被放到 artifact 里。想拿原始结构就别去 content 里抠字符串,读 artifact["structured_content"]。工具没返回结构化内容时 artifact 是 None。
服务端自己报错的情况单独说。divide 在 b=0 时抛异常,MCP 把它标成 isError=True:
1 2 type=ToolMessage status=error content=Error calling tool 'divide': division by zero
这条错误不会让 agent 崩掉,它变成一个 status="error" 的 ToolMessage 回到模型手里,模型能读到服务端自己的说法然后换个参数重试。传输层挂掉、会话断开这类故障仍然直接抛异常,因为模型对它们无能为力。
stdio 和 HTTP 的写法差异 MCPAdapter 的 target 决定了传输方式,写法上的差别只有这一个参数:
1 2 3 4 5 6 7 from pathlib import Pathfrom fastmcp import FastMCPfrom langchain.mcp import MCPAdapterin_memory = MCPAdapter(FastMCP("x" )) stdio = MCPAdapter(Path("weather_server.py" )) http = MCPAdapter("https://example.com/mcp" )
target 类型
传输
什么时候用
FastMCP 实例
进程内
单元测试,不想管进程
Path
stdio,每个 adapter 一个子进程
本地脚本,模型跑在本机
str(http/https URL)
Streamable HTTP
远程服务
MCPConfig dict
一个聚合连接,可混传输
一个 agent 接多个 server
fastmcp.Client
由你配
要自己控制 transport、缓存、认证
字符串 target 有个坑值得单独提。FastMCP 解析字符串时先当文件路径试,再当 URL 试。也就是说 "weather_server.py" 这种字符串会被当成脚本拉起来执行。配置文件和模型输出里来的 target 经常就是字符串,MCPAdapter 干脆拒绝所有不像 URL 的字符串:
1 2 3 4 ValueError: MCP target 'mcp_weather_server.py' is not a valid URL. To run a local MCP server over stdio, ask for it explicitly with Path('server.py'), a fastmcp transport, or an MCPConfig - a bare string is read as a URL so that it cannot silently become a subprocess.
本地脚本一律写 Path(...),别图省事写字符串。
HTTP 那条路我也在本机跑通了。用一个带静态 bearer token 的 server 监听 127.0.0.1:8001:
1 2 3 4 5 6 7 8 9 10 11 12 from fastmcp import FastMCPfrom fastmcp.server.auth.providers.jwt import StaticTokenVerifierauth = StaticTokenVerifier( tokens={"demo-token-123" : {"client_id" : "demo-client" , "scopes" : ["weather:read" ]}} ) mcp = FastMCP("weather-http-lab" , auth=auth) if __name__ == "__main__" : mcp.run(transport="http" , host="127.0.0.1" , port=8001 , show_banner=False )
客户端这边只换 target 和凭据:
1 2 3 4 5 6 7 8 9 10 ### 1. 带 bearer token:URL 字符串直接当 target tools: ['get_forecast'] get_forecast('Shanghai') -> {"city":"Shanghai","temp_c":19,"summary":"light rain"} ### 2. 不带 token,先看裸 HTTP 返回什么 status=401 body= ### 3. 不带 token 走 MCPAdapter MCPError: Server returned an error response
注意第 3 条。MCP 客户端把 401 包装成了 MCPError: Server returned an error response,看不到状态码。排查认证问题时先用 httpx 裸打一次,确认是 401 还是 500,别在客户端异常里猜。
多 server 并存与工具过滤 两个 server 各有一个叫 add 的工具。用 MCPConfig 把它们挂到一个 adapter 下:
1 2 3 4 5 6 7 8 9 CONFIG = { "mcpServers" : { "weather" : {"command" : sys.executable, "args" : [str (HERE / "mcp_weather_server.py" )]}, "calc" : {"command" : sys.executable, "args" : [str (HERE / "mcp_calc_server.py" )]}, } } async with MCPAdapter(CONFIG) as adapter: tools = await adapter.list_tools()
1 2 3 4 5 6 共 5 个工具: - weather_get_forecast: Return a short weather forecast for a city. - weather_add: Add two numbers. - weather_divide: Divide a by b. Fails when b is zero. - calc_add: Add two numbers (calc server version). - calc_multiply: Multiply two numbers.
FastMCP 用配置里的 key 给每个工具加了前缀,weather_add 和 calc_add 分得清清楚楚,调用也各自路由到了对的 server(两个都返回 5.0,说明没串)。config 里的 command 我写的是 sys.executable,Windows 上 "python" 经常不在 PATH 里,写死解释器路径省事。
MCPConfig 的一个限制是所有 server 共享一个协商出来的协议代次(protocol era),舰队里混进一个只支持旧握手的 server,整体会被拉到旧代次。要各走各的,用 ClientGroup,每个成员一个连接,协议代次、认证、handler 都独立:
1 2 3 4 5 6 7 8 9 10 11 from fastmcp.client import Clientfrom fastmcp.client.group import ClientGroupgroup = ClientGroup( { "weather" : Client(legacy_url, mode="legacy" ), "calc" : Client(modern_url, mode="auto" ), } ) async with MCPAdapter(group) as adapter: tools = await adapter.list_tools()
工具过滤最简单的做法是在交给 create_agent 之前筛一遍列表:
1 2 3 keep = {"weather_get_forecast" , "weather_add" } filtered = [t for t in tools if t.name in keep] agent = create_agent(model, filtered)
按注解筛比按名字筛更稳,服务端改了工具名你也不用改代码:
1 2 3 4 5 6 7 8 9 read_only = [ t for t in tools if not (t.metadata or {}) .get("mcp" , {}) .get("tool" , {}) .get("annotations" , {}) .get("destructive_hint" , False ) ]
真实输出:
1 全部 ['word_count', 'reset_counter'] -> 只读 ['word_count']
如果工具集要在运行中随状态变化,可以在 wrap_model_call 中间件里改 request.tools,那是 Tools 文档 里的动态工具选择,和 MCP 无关。工具数量一多,全塞给模型会让它挑错,我的习惯是超过二十个就做一次过滤。
连接生命周期与缓存 MCPAdapter 是异步上下文管理器,进上下文时连接,出上下文时释放。发现出来的工具自己持有 client,所以上下文退出后工具依然可调用。
flowchart LR
A[async with MCPAdapter] --> B[list_tools 读服务端目录]
B --> C[create_agent 拿到普通工具]
C --> D{每次工具调用}
D -->|默认路径| E[开连接 tools/call 关连接]
D -->|adapter 上下文仍开着| F[复用已有连接]
E --> G[ToolMessage 回到模型]
F --> G
关键在于默认路径下,一次 agent 运行里每个工具调用都会开一次会话、调用、关掉。这有两个后果。好处是长跑的 agent 不会在两次工具调用之间占着一条闲置连接。代价是同一次运行里的多次调用各建一次连接,stdio 场景下就是反复起进程。
想让一次运行共用一条会话,就把 agent 调用包在 adapter 的上下文里,重入的 client 会复用已有连接而不是再开一条。文档推荐先按默认写法来,只有在需要跨调用保持会话或者要做部署伸缩时才改。
顺带说一句源码里的事:本机这份 langchain 1.4.3 里 list_tools() 内部自己 async with self 了一次。所以外面那层 async with MCPAdapter(...) 对发现过程没有影响,它影响的是后面 agent 执行工具的那段时间。
部署场景下别每个请求重连一遍。做法是每个 run 发现一次工具目录,底下的连接池和缓存复用:
1 2 3 4 5 6 async def make_graph (): """每次 run 调一次,但复用同一个 HTTP 连接池。""" config = {"mcpServers" : {name: {"url" : url} for name, url in SERVERS.items()}} async with MCPAdapter(config) as adapter: tools = await adapter.list_tools(cache_mode="use" ) return create_agent(model, tools)
list_tools() 的 cache_mode 有三个值:
值
行为
"use"
默认。缓存里有且没过服务端给的 TTL 就用缓存,否则拉一次并写回
"refresh"
忽略缓存,重新拉并刷新缓存
"bypass"
完全跳过缓存
缓存默认不开,而且只在支持缓存提示的现代代次 server 上生效。缓存和按用户隔离是在 fastmcp.Client 上配的(Client(cache=...)),cache_mode 只决定这次发现怎么读它。
认证 凭据挂在 fastmcp.Client 上,再把 client 交给 adapter。auth 参数接受三种东西:bearer token 字符串、字面量 "oauth"、任意 httpx.Auth。
1 2 3 4 5 6 7 8 9 from fastmcp.client import Clientasync with MCPAdapter(Client(url, auth=token)) as adapter: tools = await adapter.list_tools() async with MCPAdapter(Client(url, auth="oauth" )) as adapter: tools = await adapter.list_tools()
OAuth 那条路默认把 token 存在内存里,每次运行都要重走一遍浏览器授权。要跨运行复用就自己构造一个带 token store 的 OAuth provider 传给 client,细节在 FastMCP 的 OAuth 文档 里。
静态 bearer token 和 OAuth 的差别不只是要不要跳浏览器。静态 token 由你事先签发,服务端做一次查表或验签就放行,客户端拿到就能用,中间没有任何协商。OAuth 多出三步:客户端先从 401 的 WWW-Authenticate 里找到资源元数据,读元数据拿到授权服务器地址,再动态注册自己换一个 client_id;然后走 PKCE 授权码,最后拿码换 token。多出来的这层换来的是凭据能按用户签发、能过期、能单独撤销,静态 token 一旦泄露只能整批换掉。
这些步骤不用真去接 Auth0。fastmcp 自带 InMemoryOAuthProvider,它自己就是授权服务器,签发授权码、换 token、校验 bearer 全在本地完成。下面这个 server 监听 127.0.0.1:8002:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 from fastmcp import FastMCPfrom fastmcp.server.auth.providers.in_memory import InMemoryOAuthProviderfrom fastmcp.server.dependencies import get_access_tokenfrom mcp.server.auth.settings import ClientRegistrationOptionsISSUER = "http://127.0.0.1:8002" auth = InMemoryOAuthProvider( base_url=ISSUER, client_registration_options=ClientRegistrationOptions( enabled=True , valid_scopes=["weather:read" ], ), required_scopes=["weather:read" ], ) mcp = FastMCP("weather-oauth-lab" , auth=auth) @mcp.tool(annotations={"readOnlyHint" : True } ) def get_forecast (city: str ) -> dict : """Return a short weather forecast for a city.""" ... @mcp.tool(annotations={"readOnlyHint" : True } ) def whoami () -> dict : """Report the identity the server sees for this call.""" token = get_access_token() return {"authenticated" : True , "client_id" : token.client_id, "scopes" : token.scopes}
ClientRegistrationOptions(enabled=True) 这句不能省。它默认是 False,不改的话 /register 根本不挂载,客户端做不了动态注册,auth="oauth" 会卡在注册那一步。whoami 用 get_access_token() 从当前请求里把 bearer token 对应的身份读出来,服务端到底认成了谁,一眼能看到。
手动把整条流程走一遍,每一步的原始响应如下,脚本是 MCP/OAuthClient.py :
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 ### 1. 客户端发现授权服务器 POST /mcp 不带 token -> 401 WWW-Authenticate: Bearer scope="weather:read", resource_metadata="http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp" GET http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp -> {"resource": "http://127.0.0.1:8002/mcp", "authorization_servers": ["http://127.0.0.1:8002/"], "scopes_supported": ["weather:read"], "bearer_methods_supported": ["header"]} GET http://127.0.0.1:8002/.well-known/oauth-authorization-server -> {"issuer": "http://127.0.0.1:8002/", "authorization_endpoint": "http://127.0.0.1:8002/authorize", "token_endpoint": "http://127.0.0.1:8002/token", "registration_endpoint": "http://127.0.0.1:8002/register", "scopes_supported": ["weather:read"]} ### 2. 动态客户端注册(DCR) POST http://127.0.0.1:8002/register -> 201 {"client_id": "0bd6fe0f-9651-4ac2-b36d-04a0a47f579f", "redirect_uris": ["http://127.0.0.1:8765/callback"], "scope": "weather:read", "token_endpoint_auth_method": "none"} ### 3. 授权码:GET /authorize(PKCE S256) GET /authorize -> 302 Location: http://127.0.0.1:8765/callback?code=test_auth_code_736df85f2cf490ebd2dca48304fb7051&state=XI8E7oZ-7mTGoScRsIpz4Q... code_challenge_method=S256 state 回传一致=True ### 4. 授权码换 token:POST /token POST /token -> 200 {"access_token": "test_access_token_2c0cf67a81d4733f41...", "token_type": "Bearer", "expires_in": 3600, "scope": "weather:read", "refresh_token": "test_refresh_token_1..."} ### 5. 带 token 走 MCPAdapter 调工具 tools: ['get_forecast', 'whoami'] get_forecast('Shanghai') -> {"city":"Shanghai","temp_c":19,"summary":"light rain"} whoami() -> {"authenticated":true,"client_id":"0bd6fe0f-9651-4ac2-b36d-04a0a47f579f","scopes":["weather:read"]}
第 1 步那个 401 不是故障,是入口。客户端没有 token 时先被拒一次,从响应头里拿到资源元数据地址,元数据再指向授权服务器。整条链从这一行响应头开始,你不需要预先配好任何 endpoint。InMemoryOAuthProvider 的 /authorize 直接 302 回跳,没有同意页,所以第 3 步一个 GET 就拿到了授权码。真实 IdP 会在这里停住等用户点同意,客户端代码不用改。
负面对照做了六种,都按预期失败:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 ### 6. 负面对照 --- 6a. 不带 token 裸打 /mcp status=401 WWW-Authenticate: Bearer scope="weather:read", resource_metadata="http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp" body= --- 6b. 带一个错的 token 裸打 /mcp status=401 body={"error": "invalid_token", "error_description": "Authentication failed. The provided bearer token is invalid, expired, or no longer recognized by the server. To resolve: clear authentication tokens in your MCP client and reconnect. Your client should automatically re-register and obtain new tokens."} --- 6c. 授权码复用一次(已消费的 code 再换一次 token) status=401 body={"error":"invalid_grant","error_description":"authorization code does not exist"} --- 6d. PKCE 校验失败:换 token 时给错的 code_verifier status=401 body={"error":"invalid_grant","error_description":"incorrect code_verifier"} --- 6e. 不带 token 走 MCPAdapter MCPError: Server returned an error response --- 6f. 带错 token 走 MCPAdapter MCPError: Server returned an error response
6b 和 6c 的差别值得看一眼。错 token 报 invalid_token,错授权码报 invalid_grant,都是 401,但错误码能区分是「凭据不对」还是「兑换过程不对」。6e 和 6f 又回到那个老问题:走到 MCPAdapter 里只剩一句 MCPError: Server returned an error response,401 和 500 分不出来。排查认证先裸打一次 HTTP。
还有个不用自己写这些 HTTP 调用的办法。给 Client 一个 OAuth provider,discovery、注册、PKCE、换 token 它全包了:
1 2 3 4 5 6 7 from fastmcp.client import Clientfrom fastmcp.client.auth import OAuthasync with MCPAdapter( Client(url, auth=OAuth(mcp_url=url, scopes=["weather:read" ])) ) as adapter: tools = await adapter.list_tools()
我把浏览器那一步换成一个普通 GET 实测了一遍,它自己触发一次授权跳转就拿到了 token:
1 2 3 4 5 ### 7. 交给 fastmcp.Client 自己跑:auth=OAuth provider(discovery/DCR/PKCE/token 全自动) tools: ['get_forecast', 'whoami'] get_forecast('Oslo') -> {"city":"Oslo","temp_c":3,"summary":"cloudy"} whoami() -> {"authenticated":true,"client_id":"2e0083c0-45b3-408d-b1d7-12b976576b29","scopes":["weather:read"]} 客户端自己触发了 1 次授权跳转
再提醒一句,InMemoryOAuthProvider 的 docstring 写明了是 testing provider,token 全在进程内存里,重启就没了,别拿去生产。生产接 Auth0、WorkOS 那类,客户端这半的代码一个字都不用改。
多个 server 各要各的凭据时,用 ClientGroup 给每个 client 单独设 auth:
1 2 3 4 5 6 group = ClientGroup( { "billing" : Client(billing_url, auth="oauth" ), "docs" : Client(docs_url, auth=docs_token), } )
部署到 LangGraph server 时,每次 run 应该以发起人的身份去访问 MCP server,而不是所有人共用一份凭据。分两半:认证处理器把进来的请求解析成用户身份,图工厂里按这个身份换一个 per-user token:
1 2 3 4 5 6 async def make_graph (runtime ): user = runtime.user.identity if runtime.user is not None else "anonymous" auth = BearerAuth(token_for(user)) async with MCPAdapter(Client(CONFIG, auth=auth)) as adapter: tools = await adapter.list_tools() return create_agent(model, tools)
token_for 不是某个库里的函数,它是你部署里把会话换成用户 token 的那一步:上游 OAuth 网关换好的 token,或者你自己拿 OAuth provider 按身份跑一遍上面那条授权码流程。官方文档这段就是这么写的,BearerAuth 那半能直接跑,token_for 那半得接你自己的身份系统。
这里有个我实测出来的坑。OAuth 把 token 和 client 信息写进 token store,键长这样:
1 共享 store 里的键:['mcp-oauth-token:http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-client-info:http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-token-expiry:http://127.0.0.1:8002/mcp/token_expiry']
键里只有 server URL,没有身份。两个用户共用一个 store 时,第二个用户会直接捡起第一个用户的 token,连授权跳转都不触发:
1 2 3 4 bob 和 bob2 共用一个 MemoryStore:bob 授权 1 次,bob2 授权 0 次 bob 看到的 client_id = 3e07834f-dacc-4c80-a85b-d533d906c28c bob2 看到的 client_id = 3e07834f-dacc-4c80-a85b-d533d906c28c 两者相同 -> True(bob2 直接用了 bob 的凭据,没有重新授权)
按身份分区之后就分开了。PrefixKeysWrapper 给键加前缀,两个身份共用一个底层 store 也互不干扰:
1 2 3 4 5 6 7 8 9 from key_value.aio.wrappers.prefix_keys import PrefixKeysWrapperdef oauth_for (user: str , store ) -> OAuth: """每个已验证身份一份命名空间,token 不会串到别人头上。""" return OAuth( mcp_url=DOCS_URL, scopes=["docs:read" ], token_storage=PrefixKeysWrapper(key_value=store, prefix=f"user:{user} " ), )
1 2 3 4 5 carol 和 dave 共用同一个 MemoryStore,但各自套一层 PrefixKeysWrapper:carol 授权 1 次,dave 授权 1 次 carol 看到的 client_id = 77aa045a-ff5a-47f4-9025-1cc459ebe948 dave 看到的 client_id = dee5efd9-85eb-43f5-9aab-188b4f4c6217 两者相同 -> False 分区后底层 store 里的键:['mcp-oauth-token:user:carol__http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-token:user:dave__http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-client-info:user:carol__http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-client-info:user:dave__http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-token-expiry:user:carol__http://127.0.0.1:8002/mcp/token_expiry', 'mcp-oauth-token-expiry:user:dave__http://127.0.0.1:8002/mcp/token_expiry']
工具目录的响应缓存同理。Client(cache=...) 底下那层缓存适配器要求每个存储键带一个 partition,fastmcp 的注释写着这是为了不让条目跨授权上下文碰撞,LangChain 文档也说要按验证过的身份隔离缓存;一个共享 store 的 client 舰队不分区,用户之间就会互相看到对方的目录。凭据和缓存都要挂在验证过的身份上,runtime.user.identity 是入口,后面每一步写下去的键都得带着它。
服务端反过来要能力 工具执行到一半,服务端可能需要客户端提供点东西。MCP 里有三种:elicitation(要用户填表或访问某个 URL)、sampling(要客户端帮它调一次模型)、roots(要客户端报告本地目录)。
LangChain 这边只处理 elicitation,方式是把它变成一个 LangGraph interrupt。agent 挂起,你的应用把问题展示给用户,拿到答案再恢复运行。
写服务端时有个坑我踩了。FastMCP 的 ctx.elicit() 是服务端主动推问的写法,在 2026-07-28 及之后的现代协议上没有回程通道可用,调用会直接报错:
1 elicitation via server-initiated requests is unavailable on 2026-07-28 connections.
现代代次要用 guard 模式:工具体返回一个 InputRequiredResult,客户端答完再跑一遍,服务端从 ctx.input_responses 里读答案。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 from fastmcp import Context, FastMCPfrom mcp.types import ElicitRequest, ElicitRequestFormParams, InputRequiredResultmcp = FastMCP("booking-lab" ) DATE_SCHEMA = { "type" : "object" , "properties" : {"value" : {"type" : "string" , "description" : "用餐日期" }}, "required" : ["value" ], } @mcp.tool async def book_table (people: int , ctx: Context ) -> str | InputRequiredResult: responses = ctx.input_responses if not responses: return InputRequiredResult( input_requests={ "date" : ElicitRequest( params=ElicitRequestFormParams( message="哪天来?请给一个日期。" , requested_schema=DATE_SCHEMA, ) ) } ) answer = responses["date" ] if answer.action != "accept" : return f"没订成:{answer.action} " return f"已为 {people} 人订 {answer.content['value' ]} "
客户端这一侧要给 agent 挂 checkpointer,调用和恢复用同一个 thread_id:
1 2 3 4 5 6 7 8 9 10 11 agent = create_agent(model, tools, checkpointer=InMemorySaver()) config = {"configurable" : {"thread_id" : "booking-1" }} paused = await agent.ainvoke({"messages" : [{"role" : "user" , "content" : QUESTION}]}, config) [item] = paused["__interrupt__" ] [question] = item.value["requests" ] answer = {"action" : "accept" , "content" : {"value" : "2026-09-14" }} resumed = await agent.ainvoke( Command(resume={"responses" : {question["key" ]: answer}}), config )
真实输出:
1 2 3 4 5 6 7 8 9 10 11 12 13 ### 暂停前的消息序列 [0] HumanMessage content=调用 book_table 工具,people=4。工具中途要什么你就照着流程配合,不要再反问我。 [1] AIMessage tool_call: book_table({'people': 4}) ### interrupt.value = {"type": "mcp_elicitation", "tool_name": "book_table", "requests": [{"key": "date", "message": "哪天来?请给一个日期。", "mode": "form", "requested_schema": {"properties": {"value": {"type": "string", "description": "用餐日期"}}, "required": ["value"], "type": "object"}}]} question.key=date mode=form resume 提交:{"action": "accept", "content": {"value": "2026-09-14"}} ### 恢复后的消息序列 ToolMessage status=success content=已为 4 人订 2026-09-14 最后一条:已为您订好:4 人,2026-09-14。
stateDiagram-v2
[*] --> 调用工具
调用工具 --> 等待人工输入: 服务端返回 InputRequiredResult
等待人工输入 --> 重跑工具: Command resume responses
重跑工具 --> 返回结果: 服务端不再要输入
重跑工具 --> 等待人工输入: 又一轮提问
等待人工输入 --> 放弃: action 为 decline 或 cancel
答案的 key 必须和服务端给的 key 对上,少一个都会报错。action 只有三个值:accept(带符合 schema 的 content)、decline、cancel。
恢复会把工具从头再跑一遍。服务端在提问之前做过的活儿会重做,写工具时要保证重复执行不出副作用。我那个 book_table 是先提问再订座,所以没有重复下单的问题;如果顺序反了,就要自己做幂等。
sampling 和 roots 没有这条路。MCP 工具文档 和迁移说明 把话说死了:langchain.mcp 不答 sampling(服务端要客户端帮它跑一次模型),也不答 roots(服务端问客户端能碰到哪些本地路径),工具调用返回这两类请求会抛 NotImplementedError。
这条我没实测,原因值得单独说清楚。要触发那个 NotImplementedError,得有一个真会向客户端发起 sampling 或 roots 请求的 server;而 FastMCP 4 把 ctx.sample() 和 ctx.list_roots() 从所有协议世代里删掉了。发起端没有口子,客户端那条错误分支就走不到。所以这不是”没来得及跑”,是当前工具链下构造不出这个场景,我只能引文档。
迁移说明里还有句容易误读的:老包用 Callbacks 挂 handler,那个对象在 langchain.mcp 里已经没了,但底层 handler 还在,FastMCP 的 Client 直接收,例如 progress_handler、log_handler。那张对照表里没有 sampling 和 roots 的 handler,官方给的路子是开 issue。所以”自己挂个 handler 就能答 sampling”是不成立的。elicitation 是另一回事:它已经被 adapter 自动接成 interrupt,只有当你传进去的 client 自带 elicitation_handler 时,adapter 才会尊重你的那个、不覆盖。
没有 server 时怎么本地自测 三种办法,从轻到重。
进程内 FastMCP 实例最省事,没有子进程也没有 socket,适合写成单元测试:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 from fastmcp import FastMCPfrom langchain.mcp import MCPAdaptermcp = FastMCP("in-memory-lab" ) @mcp.tool(annotations={"readOnlyHint" : True } ) def word_count (text: str ) -> int : """Count whitespace-separated words.""" return len (text.split()) async def main (): async with MCPAdapter(mcp) as adapter: tools = await adapter.list_tools() msg = await tools[0 ].ainvoke( {"type" : "tool_call" , "name" : "word_count" , "args" : {"text" : "a b c d" }, "id" : "c1" } ) print (msg.content[0 ]["text" ])
1 2 3 4 5 6 7 8 9 ### 进程内 FastMCP server:没有子进程,没有 socket - word_count: Count whitespace-separated words. - reset_counter: Pretend to wipe a counter. ### 按 destructive_hint 过滤,只留只读工具 全部 ['word_count', 'reset_counter'] -> 只读 ['word_count'] ### 直接调用 word_count('a b c d') -> 4
第二种是 Path 加一个最小脚本,也就是前面 stdio 那段,验证的是真实的进程拉起和协议握手。
第三种是接上模型跑整条链。这套组合(本地 stdio server + DeepSeek + create_agent)在本机是通的:
1 2 3 4 async with MCPAdapter(SERVER) as adapter: tools = await adapter.list_tools() agent = create_agent(model, tools) result = await agent.ainvoke({"messages" : [{"role" : "user" , "content" : QUESTION}]})
下面这段做了截断,只留能说明问题的几行:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 ### 一次 agent 运行的完整消息序列 [0] HumanMessage content: 先查 Oslo 的天气,再算 10 除以 0。两次工具返回的原文都要贴出来。 [1] AIMessage tool_call: get_forecast({'city': 'Oslo'}) tool_call: divide({'a': 10, 'b': 0}) [2] ToolMessage status: success content: {"city":"Oslo","temp_c":3,"summary":"cloudy"} artifact: {'structured_content': {'city': 'Oslo', 'temp_c': 3, 'summary': 'cloudy'}} [3] ToolMessage status: error content: Error calling tool 'divide': division by zero [4] AIMessage content: 两次调用的原文返回如下:1. Oslo 天气 get_forecast: {"city":"Oslo","temp_c":3,"summary":"cloudy"} 2. 10 除以 0 divide:Error calling tool 'divide': division by zero
模型一口气发了两个 tool_call,成功的和失败的各回一条 ToolMessage,它照着两份原文组织答案。到这一步,MCP 这条链在你的机器上就算通了,换远程 server 只是把 Path 换成 URL 加凭据。
小结
MCP 把工具的声明和调用协议化了。接入成本从「每个数据源写一份适配」降到「换一个 target」。
LangChain 这边的入口是 langchain.mcp.MCPAdapter,list_tools() 拿到的就是普通 LangChain 工具,可以直接进 create_agent。它还是 beta,记得锁版本。
传输由 target 的类型决定:Path 走 stdio,http/https 字符串走 Streamable HTTP,FastMCP 实例走进程内。本地脚本一律用 Path,字符串会被当成 URL 校验。
默认每个工具调用开一次会话。要跨调用复用就把 agent 调用包在 adapter 上下文里;部署时按 run 发现工具目录,底下用连接池和 cache_mode 扛开销。
服务端要输入只有 elicitation 能答,落到 LangGraph interrupt。服务端要用 guard 模式返回 InputRequiredResult,且恢复会重跑工具,副作用要做幂等。