代码仓库ChainReaction

给 agent 接工具,真正花时间的部分往往不是写函数体。函数体十行就写完了,剩下的时间花在接口对齐上:这个数据源要什么参数,那个数据源返回的结构怎么塞给模型,调用失败时错误信息用什么格式回传。每接一个系统,这套活儿重来一遍。

MCP(Model Context Protocol)把这层约定协议化了。服务端自己声明它有哪些工具、每个工具叫什么、参数 schema 长什么样;客户端连上去把这些声明读出来,转成模型能看懂的格式。适配代码从「每个数据源写一遍」变成「一份通用实现」。

LangChain 从 1.4 开始把 MCP 客户端收进了主包,命名空间是 langchain.mcp,核心类叫 MCPAdapter。它在 FastMCP 之上做转换:发现服务端的工具,变成普通的 LangChain 工具,直接丢给 create_agent。传输、协议握手、认证这些往下的事由 FastMCP 负责。

这篇按能跑的顺序讲:调用链长什么样,stdio 和 HTTP 两种连法差在哪,多个 server 挤在一起时工具名怎么处理,连接什么时候开什么时候关,认证怎么挂,服务端反过来要客户端能力时怎么办,手边一个 MCP server 都没有时怎么本地自测。文中的输出都是在本机跑出来的,脚本放在 MCP/。

MCP 标准化了哪一层

一个 MCP server 对外暴露的东西很薄:工具名、描述、输入 JSON Schema,外加一组可选的注解(比如 readOnlyHint、destructiveHint)。客户端用 tools/list 把这份目录读回来,用 tools/call 执行,返回值里带内容块、可选的 structuredContent,以及一个 isError 标志。

就这么几样东西,但它把「怎么发现工具」和「怎么调用工具」固定下来了。以前你为 GitHub 写一个 search_repos 工具,为内部知识库写一个 search_docs 工具,两份代码里的参数校验、错误处理、返回值包装各写各的。现在只要对方提供了 MCP server,你换一个 target 就完事,agent 那边的代码一行不动。

代价是引入了一层间接。工具的 schema 由服务端定义,你不再能随手改描述文案;服务端挂了,你的 agent 就少一块能力。这笔账划不划算,取决于你要接多少个外部系统。接三个以内,自己写工具更快;接十个,MCP 省下来的维护成本就很明显了。

调用链与接入点

从 agent 的角度看,MCP 工具和 @tool 装饰出来的工具没有区别。多出来的两层是 MCPAdapter 做的转换和 FastMCP 做的传输。

装依赖时注意,MCP 是主包的额外依赖:

1
pip install "langchain[mcp]"

这个命名空间还处于 beta,导入时会抛一次 LangChainBetaWarning。API 可能变,生产环境锁版本。

最小用法是三步:开 adapter、列工具、建 agent。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
import asyncio
import os

from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI

model = ChatOpenAI(
api_key=os.getenv("DEEPSEEK_API_KEY"),
base_url="https://api.deepseek.com/v1",
model="deepseek-chat",
temperature=0.1,
max_tokens=1000,
)


async def main():
async with MCPAdapter("https://example.com/mcp") as adapter:
tools = await adapter.list_tools()
agent = create_agent(model, tools)
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": "查一下 Oslo 的天气"}]}
)
return result


asyncio.run(main())

没有现成 server 也能验证这条链。下面这个 server 只有三十行,用 FastMCP 写,跑 stdio。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
from fastmcp import FastMCP

mcp = FastMCP("weather-lab", instructions="演示用天气 / 算术 MCP server")

_FORECAST = {
"Oslo": {"temp_c": 3, "summary": "cloudy"},
"Shanghai": {"temp_c": 19, "summary": "light rain"},
}


@mcp.tool(annotations={"readOnlyHint": True, "title": "查天气"})
def get_forecast(city: str) -> dict:
"""Return a short weather forecast for a city."""
data = _FORECAST.get(city, {"temp_c": 20, "summary": "unknown city, fallback"})
return {"city": city, **data}


@mcp.tool(annotations={"readOnlyHint": True})
def add(a: float, b: float) -> float:
"""Add two numbers."""
return a + b


@mcp.tool(annotations={"destructiveHint": True, "readOnlyHint": False})
def divide(a: float, b: float) -> float:
"""Divide a by b. Fails when b is zero."""
if b == 0:
raise ValueError("division by zero")
return a / b


if __name__ == "__main__":
mcp.run(transport="stdio", show_banner=False)

连上去列一遍工具,看看适配结果长什么样。我用 Path 指脚本,走 stdio:

1
2
3
4
SERVER = Path(__file__).with_name("mcp_weather_server.py")

async with MCPAdapter(SERVER) as adapter:
tools = await adapter.list_tools()

真实输出(省略了 add 那条,结构和 get_forecast 一样):

1
2
3
4
5
6
- get_forecast: Return a short weather forecast for a city.
args_schema: {"type": "object", "additionalProperties": false, "properties": {"city": {"type": "string"}}, "required": ["city"]}
metadata: {"mcp": {"tool": {"annotations": {"title": "查天气", "read_only_hint": true}, "_meta": {"fastmcp": {"tags": []}}}, "server": {"name": "weather-lab", "version": "4.0.10"}}}
- divide: Divide a by b. Fails when b is zero.
args_schema: {"type": "object", "additionalProperties": false, "properties": {"a": {"type": "number"}, "b": {"type": "number"}}, "required": ["a", "b"]}
metadata: {"mcp": {"tool": {"annotations": {"read_only_hint": false, "destructive_hint": true}, ...}, "server": {"name": "weather-lab", "version": "4.0.10"}}}

服务端写的注解原样落到了 LangChain 工具的 metadata["mcp"] 下面,readOnlyHint 变成 snake_case 的 read_only_hint,服务端身份在 mcp.server 里。这些字段全是可选的,读的时候要带默认值,不能假设一定存在。

不经过模型直接调工具,能看清返回值被拆成了什么。传一个完整的 ToolCall 进去:

1
2
3
msg = await forecast.ainvoke(
{"type": "tool_call", "name": "get_forecast", "args": {"city": "Oslo"}, "id": "call_1"}
)
1
2
3
type=ToolMessage status=success
content={"city":"Oslo","temp_c":3,"summary":"cloudy"}
artifact={'structured_content': {'city': 'Oslo', 'temp_c': 3, 'summary': 'cloudy'}}

工具返回 dict,模型看到的是压成一行 JSON 的文本,结构化的那份被放到 artifact 里。想拿原始结构就别去 content 里抠字符串,读 artifact["structured_content"]。工具没返回结构化内容时 artifact 是 None。

服务端自己报错的情况单独说。divide 在 b=0 时抛异常,MCP 把它标成 isError=True:

1
2
type=ToolMessage status=error
content=Error calling tool 'divide': division by zero

这条错误不会让 agent 崩掉,它变成一个 status="error" 的 ToolMessage 回到模型手里,模型能读到服务端自己的说法然后换个参数重试。传输层挂掉、会话断开这类故障仍然直接抛异常,因为模型对它们无能为力。

stdio 和 HTTP 的写法差异

MCPAdapter 的 target 决定了传输方式,写法上的差别只有这一个参数:

1
2
3
4
5
6
7
from pathlib import Path
from fastmcp import FastMCP
from langchain.mcp import MCPAdapter

in_memory = MCPAdapter(FastMCP("x")) # 进程内,无子进程无 socket
stdio = MCPAdapter(Path("weather_server.py")) # 拉子进程,走 stdio
http = MCPAdapter("https://example.com/mcp") # Streamable HTTP
target 类型 传输 什么时候用
FastMCP 实例 进程内 单元测试,不想管进程
Path stdio,每个 adapter 一个子进程 本地脚本,模型跑在本机
str(http/https URL) Streamable HTTP 远程服务
MCPConfig dict 一个聚合连接,可混传输 一个 agent 接多个 server
fastmcp.Client 由你配 要自己控制 transport、缓存、认证

字符串 target 有个坑值得单独提。FastMCP 解析字符串时先当文件路径试,再当 URL 试。也就是说 "weather_server.py" 这种字符串会被当成脚本拉起来执行。配置文件和模型输出里来的 target 经常就是字符串,MCPAdapter 干脆拒绝所有不像 URL 的字符串:

1
2
3
4
ValueError: MCP target 'mcp_weather_server.py' is not a valid URL. To run a local MCP
server over stdio, ask for it explicitly with Path('server.py'), a fastmcp transport,
or an MCPConfig - a bare string is read as a URL so that it cannot silently become a
subprocess.

本地脚本一律写 Path(...),别图省事写字符串。

HTTP 那条路我也在本机跑通了。用一个带静态 bearer token 的 server 监听 127.0.0.1:8001:

1
2
3
4
5
6
7
8
9
10
11
12
from fastmcp import FastMCP
from fastmcp.server.auth.providers.jwt import StaticTokenVerifier

auth = StaticTokenVerifier(
tokens={"demo-token-123": {"client_id": "demo-client", "scopes": ["weather:read"]}}
)
mcp = FastMCP("weather-http-lab", auth=auth)

# ... 工具定义和 stdio 版本一样 ...

if __name__ == "__main__":
mcp.run(transport="http", host="127.0.0.1", port=8001, show_banner=False)

客户端这边只换 target 和凭据:

1
2
3
4
5
6
7
8
9
10
### 1. 带 bearer token:URL 字符串直接当 target
tools: ['get_forecast']
get_forecast('Shanghai') -> {"city":"Shanghai","temp_c":19,"summary":"light rain"}

### 2. 不带 token,先看裸 HTTP 返回什么
status=401
body=

### 3. 不带 token 走 MCPAdapter
MCPError: Server returned an error response

注意第 3 条。MCP 客户端把 401 包装成了 MCPError: Server returned an error response,看不到状态码。排查认证问题时先用 httpx 裸打一次,确认是 401 还是 500,别在客户端异常里猜。

多 server 并存与工具过滤

两个 server 各有一个叫 add 的工具。用 MCPConfig 把它们挂到一个 adapter 下:

1
2
3
4
5
6
7
8
9
CONFIG = {
"mcpServers": {
"weather": {"command": sys.executable, "args": [str(HERE / "mcp_weather_server.py")]},
"calc": {"command": sys.executable, "args": [str(HERE / "mcp_calc_server.py")]},
}
}

async with MCPAdapter(CONFIG) as adapter:
tools = await adapter.list_tools()
1
2
3
4
5
6
共 5 个工具:
- weather_get_forecast: Return a short weather forecast for a city.
- weather_add: Add two numbers.
- weather_divide: Divide a by b. Fails when b is zero.
- calc_add: Add two numbers (calc server version).
- calc_multiply: Multiply two numbers.

FastMCP 用配置里的 key 给每个工具加了前缀,weather_add 和 calc_add 分得清清楚楚,调用也各自路由到了对的 server(两个都返回 5.0,说明没串)。config 里的 command 我写的是 sys.executable,Windows 上 "python" 经常不在 PATH 里,写死解释器路径省事。

MCPConfig 的一个限制是所有 server 共享一个协商出来的协议代次(protocol era),舰队里混进一个只支持旧握手的 server,整体会被拉到旧代次。要各走各的,用 ClientGroup,每个成员一个连接,协议代次、认证、handler 都独立:

1
2
3
4
5
6
7
8
9
10
11
from fastmcp.client import Client
from fastmcp.client.group import ClientGroup

group = ClientGroup(
{
"weather": Client(legacy_url, mode="legacy"),
"calc": Client(modern_url, mode="auto"),
}
)
async with MCPAdapter(group) as adapter:
tools = await adapter.list_tools()

工具过滤最简单的做法是在交给 create_agent 之前筛一遍列表:

1
2
3
keep = {"weather_get_forecast", "weather_add"}
filtered = [t for t in tools if t.name in keep]
agent = create_agent(model, filtered)

按注解筛比按名字筛更稳,服务端改了工具名你也不用改代码:

1
2
3
4
5
6
7
8
9
read_only = [
t
for t in tools
if not (t.metadata or {})
.get("mcp", {})
.get("tool", {})
.get("annotations", {})
.get("destructive_hint", False)
]

真实输出:

1
全部 ['word_count', 'reset_counter'] -> 只读 ['word_count']

如果工具集要在运行中随状态变化,可以在 wrap_model_call 中间件里改 request.tools,那是 Tools 文档 里的动态工具选择,和 MCP 无关。工具数量一多,全塞给模型会让它挑错,我的习惯是超过二十个就做一次过滤。

连接生命周期与缓存

MCPAdapter 是异步上下文管理器,进上下文时连接,出上下文时释放。发现出来的工具自己持有 client,所以上下文退出后工具依然可调用。

关键在于默认路径下,一次 agent 运行里每个工具调用都会开一次会话、调用、关掉。这有两个后果。好处是长跑的 agent 不会在两次工具调用之间占着一条闲置连接。代价是同一次运行里的多次调用各建一次连接,stdio 场景下就是反复起进程。

想让一次运行共用一条会话,就把 agent 调用包在 adapter 的上下文里,重入的 client 会复用已有连接而不是再开一条。文档推荐先按默认写法来,只有在需要跨调用保持会话或者要做部署伸缩时才改。

顺带说一句源码里的事:本机这份 langchain 1.4.3 里 list_tools() 内部自己 async with self 了一次。所以外面那层 async with MCPAdapter(...) 对发现过程没有影响,它影响的是后面 agent 执行工具的那段时间。

部署场景下别每个请求重连一遍。做法是每个 run 发现一次工具目录,底下的连接池和缓存复用:

1
2
3
4
5
6
async def make_graph():
"""每次 run 调一次,但复用同一个 HTTP 连接池。"""
config = {"mcpServers": {name: {"url": url} for name, url in SERVERS.items()}}
async with MCPAdapter(config) as adapter:
tools = await adapter.list_tools(cache_mode="use")
return create_agent(model, tools)

list_tools() 的 cache_mode 有三个值:

值 行为
"use" 默认。缓存里有且没过服务端给的 TTL 就用缓存,否则拉一次并写回
"refresh" 忽略缓存,重新拉并刷新缓存
"bypass" 完全跳过缓存

缓存默认不开,而且只在支持缓存提示的现代代次 server 上生效。缓存和按用户隔离是在 fastmcp.Client 上配的(Client(cache=...)),cache_mode 只决定这次发现怎么读它。

认证

凭据挂在 fastmcp.Client 上,再把 client 交给 adapter。auth 参数接受三种东西:bearer token 字符串、字面量 "oauth"、任意 httpx.Auth。

1
2
3
4
5
6
7
8
9
from fastmcp.client import Client

# 静态 token,没有 discovery,没有浏览器跳转
async with MCPAdapter(Client(url, auth=token)) as adapter:
tools = await adapter.list_tools()

# 完整的 OAuth 2.1:discovery、动态客户端注册、浏览器跳转、换 token
async with MCPAdapter(Client(url, auth="oauth")) as adapter:
tools = await adapter.list_tools()

OAuth 那条路默认把 token 存在内存里,每次运行都要重走一遍浏览器授权。要跨运行复用就自己构造一个带 token store 的 OAuth provider 传给 client,细节在 FastMCP 的 OAuth 文档 里。

静态 bearer token 和 OAuth 的差别不只是要不要跳浏览器。静态 token 由你事先签发,服务端做一次查表或验签就放行,客户端拿到就能用,中间没有任何协商。OAuth 多出三步:客户端先从 401 的 WWW-Authenticate 里找到资源元数据,读元数据拿到授权服务器地址,再动态注册自己换一个 client_id;然后走 PKCE 授权码,最后拿码换 token。多出来的这层换来的是凭据能按用户签发、能过期、能单独撤销,静态 token 一旦泄露只能整批换掉。

这些步骤不用真去接 Auth0。fastmcp 自带 InMemoryOAuthProvider,它自己就是授权服务器,签发授权码、换 token、校验 bearer 全在本地完成。下面这个 server 监听 127.0.0.1:8002:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
from fastmcp import FastMCP
from fastmcp.server.auth.providers.in_memory import InMemoryOAuthProvider
from fastmcp.server.dependencies import get_access_token
from mcp.server.auth.settings import ClientRegistrationOptions

ISSUER = "http://127.0.0.1:8002"

auth = InMemoryOAuthProvider(
base_url=ISSUER,
client_registration_options=ClientRegistrationOptions(
enabled=True,
valid_scopes=["weather:read"],
),
required_scopes=["weather:read"],
)

mcp = FastMCP("weather-oauth-lab", auth=auth)


@mcp.tool(annotations={"readOnlyHint": True})
def get_forecast(city: str) -> dict:
"""Return a short weather forecast for a city."""
...


@mcp.tool(annotations={"readOnlyHint": True})
def whoami() -> dict:
"""Report the identity the server sees for this call."""
token = get_access_token()
return {"authenticated": True, "client_id": token.client_id, "scopes": token.scopes}

ClientRegistrationOptions(enabled=True) 这句不能省。它默认是 False,不改的话 /register 根本不挂载,客户端做不了动态注册,auth="oauth" 会卡在注册那一步。whoami 用 get_access_token() 从当前请求里把 bearer token 对应的身份读出来,服务端到底认成了谁,一眼能看到。

手动把整条流程走一遍,每一步的原始响应如下,脚本是 MCP/OAuthClient.py:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
### 1. 客户端发现授权服务器
POST /mcp 不带 token -> 401
WWW-Authenticate: Bearer scope="weather:read", resource_metadata="http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp"
GET http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp -> {"resource": "http://127.0.0.1:8002/mcp", "authorization_servers": ["http://127.0.0.1:8002/"], "scopes_supported": ["weather:read"], "bearer_methods_supported": ["header"]}
GET http://127.0.0.1:8002/.well-known/oauth-authorization-server ->
{"issuer": "http://127.0.0.1:8002/", "authorization_endpoint": "http://127.0.0.1:8002/authorize", "token_endpoint": "http://127.0.0.1:8002/token", "registration_endpoint": "http://127.0.0.1:8002/register", "scopes_supported": ["weather:read"]}

### 2. 动态客户端注册(DCR)
POST http://127.0.0.1:8002/register -> 201
{"client_id": "0bd6fe0f-9651-4ac2-b36d-04a0a47f579f", "redirect_uris": ["http://127.0.0.1:8765/callback"], "scope": "weather:read", "token_endpoint_auth_method": "none"}

### 3. 授权码:GET /authorize(PKCE S256)
GET /authorize -> 302
Location: http://127.0.0.1:8765/callback?code=test_auth_code_736df85f2cf490ebd2dca48304fb7051&state=XI8E7oZ-7mTGoScRsIpz4Q...
code_challenge_method=S256 state 回传一致=True

### 4. 授权码换 token:POST /token
POST /token -> 200
{"access_token": "test_access_token_2c0cf67a81d4733f41...", "token_type": "Bearer", "expires_in": 3600, "scope": "weather:read", "refresh_token": "test_refresh_token_1..."}

### 5. 带 token 走 MCPAdapter 调工具
tools: ['get_forecast', 'whoami']
get_forecast('Shanghai') -> {"city":"Shanghai","temp_c":19,"summary":"light rain"}
whoami() -> {"authenticated":true,"client_id":"0bd6fe0f-9651-4ac2-b36d-04a0a47f579f","scopes":["weather:read"]}

第 1 步那个 401 不是故障,是入口。客户端没有 token 时先被拒一次,从响应头里拿到资源元数据地址,元数据再指向授权服务器。整条链从这一行响应头开始,你不需要预先配好任何 endpoint。InMemoryOAuthProvider 的 /authorize 直接 302 回跳,没有同意页,所以第 3 步一个 GET 就拿到了授权码。真实 IdP 会在这里停住等用户点同意,客户端代码不用改。

负面对照做了六种,都按预期失败:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
### 6. 负面对照
--- 6a. 不带 token 裸打 /mcp
status=401
WWW-Authenticate: Bearer scope="weather:read", resource_metadata="http://127.0.0.1:8002/.well-known/oauth-protected-resource/mcp"
body=
--- 6b. 带一个错的 token 裸打 /mcp
status=401
body={"error": "invalid_token", "error_description": "Authentication failed. The provided bearer token is invalid, expired, or no longer recognized by the server. To resolve: clear authentication tokens in your MCP client and reconnect. Your client should automatically re-register and obtain new tokens."}
--- 6c. 授权码复用一次(已消费的 code 再换一次 token)
status=401 body={"error":"invalid_grant","error_description":"authorization code does not exist"}
--- 6d. PKCE 校验失败:换 token 时给错的 code_verifier
status=401 body={"error":"invalid_grant","error_description":"incorrect code_verifier"}
--- 6e. 不带 token 走 MCPAdapter
MCPError: Server returned an error response
--- 6f. 带错 token 走 MCPAdapter
MCPError: Server returned an error response

6b 和 6c 的差别值得看一眼。错 token 报 invalid_token,错授权码报 invalid_grant,都是 401,但错误码能区分是「凭据不对」还是「兑换过程不对」。6e 和 6f 又回到那个老问题:走到 MCPAdapter 里只剩一句 MCPError: Server returned an error response,401 和 500 分不出来。排查认证先裸打一次 HTTP。

还有个不用自己写这些 HTTP 调用的办法。给 Client 一个 OAuth provider,discovery、注册、PKCE、换 token 它全包了:

1
2
3
4
5
6
7
from fastmcp.client import Client
from fastmcp.client.auth import OAuth

async with MCPAdapter(
Client(url, auth=OAuth(mcp_url=url, scopes=["weather:read"]))
) as adapter:
tools = await adapter.list_tools()

我把浏览器那一步换成一个普通 GET 实测了一遍,它自己触发一次授权跳转就拿到了 token:

1
2
3
4
5
### 7. 交给 fastmcp.Client 自己跑:auth=OAuth provider(discovery/DCR/PKCE/token 全自动)
tools: ['get_forecast', 'whoami']
get_forecast('Oslo') -> {"city":"Oslo","temp_c":3,"summary":"cloudy"}
whoami() -> {"authenticated":true,"client_id":"2e0083c0-45b3-408d-b1d7-12b976576b29","scopes":["weather:read"]}
客户端自己触发了 1 次授权跳转

再提醒一句,InMemoryOAuthProvider 的 docstring 写明了是 testing provider,token 全在进程内存里,重启就没了,别拿去生产。生产接 Auth0、WorkOS 那类,客户端这半的代码一个字都不用改。

多个 server 各要各的凭据时,用 ClientGroup 给每个 client 单独设 auth:

1
2
3
4
5
6
group = ClientGroup(
{
"billing": Client(billing_url, auth="oauth"),
"docs": Client(docs_url, auth=docs_token),
}
)

部署到 LangGraph server 时,每次 run 应该以发起人的身份去访问 MCP server,而不是所有人共用一份凭据。分两半:认证处理器把进来的请求解析成用户身份,图工厂里按这个身份换一个 per-user token:

1
2
3
4
5
6
async def make_graph(runtime):
user = runtime.user.identity if runtime.user is not None else "anonymous"
auth = BearerAuth(token_for(user)) # 换成这个用户的 token
async with MCPAdapter(Client(CONFIG, auth=auth)) as adapter:
tools = await adapter.list_tools()
return create_agent(model, tools)

token_for 不是某个库里的函数,它是你部署里把会话换成用户 token 的那一步:上游 OAuth 网关换好的 token,或者你自己拿 OAuth provider 按身份跑一遍上面那条授权码流程。官方文档这段就是这么写的,BearerAuth 那半能直接跑,token_for 那半得接你自己的身份系统。

这里有个我实测出来的坑。OAuth 把 token 和 client 信息写进 token store,键长这样:

1
共享 store 里的键:['mcp-oauth-token:http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-client-info:http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-token-expiry:http://127.0.0.1:8002/mcp/token_expiry']

键里只有 server URL,没有身份。两个用户共用一个 store 时,第二个用户会直接捡起第一个用户的 token,连授权跳转都不触发:

1
2
3
4
bob 和 bob2 共用一个 MemoryStore:bob 授权 1 次,bob2 授权 0 次
bob 看到的 client_id = 3e07834f-dacc-4c80-a85b-d533d906c28c
bob2 看到的 client_id = 3e07834f-dacc-4c80-a85b-d533d906c28c
两者相同 -> True(bob2 直接用了 bob 的凭据,没有重新授权)

按身份分区之后就分开了。PrefixKeysWrapper 给键加前缀,两个身份共用一个底层 store 也互不干扰:

1
2
3
4
5
6
7
8
9
from key_value.aio.wrappers.prefix_keys import PrefixKeysWrapper

def oauth_for(user: str, store) -> OAuth:
"""每个已验证身份一份命名空间,token 不会串到别人头上。"""
return OAuth(
mcp_url=DOCS_URL,
scopes=["docs:read"],
token_storage=PrefixKeysWrapper(key_value=store, prefix=f"user:{user}"),
)
1
2
3
4
5
carol 和 dave 共用同一个 MemoryStore,但各自套一层 PrefixKeysWrapper:carol 授权 1 次,dave 授权 1 次
carol 看到的 client_id = 77aa045a-ff5a-47f4-9025-1cc459ebe948
dave 看到的 client_id = dee5efd9-85eb-43f5-9aab-188b4f4c6217
两者相同 -> False
分区后底层 store 里的键:['mcp-oauth-token:user:carol__http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-token:user:dave__http://127.0.0.1:8002/mcp/tokens', 'mcp-oauth-client-info:user:carol__http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-client-info:user:dave__http://127.0.0.1:8002/mcp/client_info', 'mcp-oauth-token-expiry:user:carol__http://127.0.0.1:8002/mcp/token_expiry', 'mcp-oauth-token-expiry:user:dave__http://127.0.0.1:8002/mcp/token_expiry']

工具目录的响应缓存同理。Client(cache=...) 底下那层缓存适配器要求每个存储键带一个 partition,fastmcp 的注释写着这是为了不让条目跨授权上下文碰撞,LangChain 文档也说要按验证过的身份隔离缓存;一个共享 store 的 client 舰队不分区,用户之间就会互相看到对方的目录。凭据和缓存都要挂在验证过的身份上,runtime.user.identity 是入口,后面每一步写下去的键都得带着它。

服务端反过来要能力

工具执行到一半,服务端可能需要客户端提供点东西。MCP 里有三种:elicitation(要用户填表或访问某个 URL)、sampling(要客户端帮它调一次模型)、roots(要客户端报告本地目录)。

LangChain 这边只处理 elicitation,方式是把它变成一个 LangGraph interrupt。agent 挂起,你的应用把问题展示给用户,拿到答案再恢复运行。

写服务端时有个坑我踩了。FastMCP 的 ctx.elicit() 是服务端主动推问的写法,在 2026-07-28 及之后的现代协议上没有回程通道可用,调用会直接报错:

1
elicitation via server-initiated requests is unavailable on 2026-07-28 connections.

现代代次要用 guard 模式:工具体返回一个 InputRequiredResult,客户端答完再跑一遍,服务端从 ctx.input_responses 里读答案。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
from fastmcp import Context, FastMCP
from mcp.types import ElicitRequest, ElicitRequestFormParams, InputRequiredResult

mcp = FastMCP("booking-lab")

DATE_SCHEMA = {
"type": "object",
"properties": {"value": {"type": "string", "description": "用餐日期"}},
"required": ["value"],
}


@mcp.tool
async def book_table(people: int, ctx: Context) -> str | InputRequiredResult:
responses = ctx.input_responses
if not responses:
return InputRequiredResult(
input_requests={
"date": ElicitRequest(
params=ElicitRequestFormParams(
message="哪天来?请给一个日期。",
requested_schema=DATE_SCHEMA,
)
)
}
)
answer = responses["date"]
if answer.action != "accept":
return f"没订成:{answer.action}"
return f"已为 {people} 人订 {answer.content['value']}"

客户端这一侧要给 agent 挂 checkpointer,调用和恢复用同一个 thread_id:

1
2
3
4
5
6
7
8
9
10
11
agent = create_agent(model, tools, checkpointer=InMemorySaver())
config = {"configurable": {"thread_id": "booking-1"}}

paused = await agent.ainvoke({"messages": [{"role": "user", "content": QUESTION}]}, config)
[item] = paused["__interrupt__"]
[question] = item.value["requests"]

answer = {"action": "accept", "content": {"value": "2026-09-14"}}
resumed = await agent.ainvoke(
Command(resume={"responses": {question["key"]: answer}}), config
)

真实输出:

1
2
3
4
5
6
7
8
9
10
11
12
13
### 暂停前的消息序列
[0] HumanMessage
content=调用 book_table 工具,people=4。工具中途要什么你就照着流程配合,不要再反问我。
[1] AIMessage
tool_call: book_table({'people': 4})

### interrupt.value = {"type": "mcp_elicitation", "tool_name": "book_table", "requests": [{"key": "date", "message": "哪天来?请给一个日期。", "mode": "form", "requested_schema": {"properties": {"value": {"type": "string", "description": "用餐日期"}}, "required": ["value"], "type": "object"}}]}
question.key=date mode=form
resume 提交:{"action": "accept", "content": {"value": "2026-09-14"}}

### 恢复后的消息序列
ToolMessage status=success content=已为 4 人订 2026-09-14
最后一条:已为您订好:4 人,2026-09-14。

答案的 key 必须和服务端给的 key 对上,少一个都会报错。action 只有三个值:accept(带符合 schema 的 content)、decline、cancel。

恢复会把工具从头再跑一遍。服务端在提问之前做过的活儿会重做,写工具时要保证重复执行不出副作用。我那个 book_table 是先提问再订座,所以没有重复下单的问题;如果顺序反了,就要自己做幂等。

sampling 和 roots 没有这条路。MCP 工具文档 和迁移说明把话说死了:langchain.mcp 不答 sampling(服务端要客户端帮它跑一次模型),也不答 roots(服务端问客户端能碰到哪些本地路径),工具调用返回这两类请求会抛 NotImplementedError。

这条我没实测,原因值得单独说清楚。要触发那个 NotImplementedError,得有一个真会向客户端发起 sampling 或 roots 请求的 server;而 FastMCP 4 把 ctx.sample() 和 ctx.list_roots() 从所有协议世代里删掉了。发起端没有口子,客户端那条错误分支就走不到。所以这不是”没来得及跑”,是当前工具链下构造不出这个场景,我只能引文档。

迁移说明里还有句容易误读的:老包用 Callbacks 挂 handler,那个对象在 langchain.mcp 里已经没了,但底层 handler 还在,FastMCP 的 Client 直接收,例如 progress_handler、log_handler。那张对照表里没有 sampling 和 roots 的 handler,官方给的路子是开 issue。所以”自己挂个 handler 就能答 sampling”是不成立的。elicitation 是另一回事:它已经被 adapter 自动接成 interrupt,只有当你传进去的 client 自带 elicitation_handler 时,adapter 才会尊重你的那个、不覆盖。

没有 server 时怎么本地自测

三种办法,从轻到重。

进程内 FastMCP 实例最省事,没有子进程也没有 socket,适合写成单元测试:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from fastmcp import FastMCP
from langchain.mcp import MCPAdapter

mcp = FastMCP("in-memory-lab")


@mcp.tool(annotations={"readOnlyHint": True})
def word_count(text: str) -> int:
"""Count whitespace-separated words."""
return len(text.split())


async def main():
async with MCPAdapter(mcp) as adapter:
tools = await adapter.list_tools()
msg = await tools[0].ainvoke(
{"type": "tool_call", "name": "word_count", "args": {"text": "a b c d"}, "id": "c1"}
)
print(msg.content[0]["text"])
1
2
3
4
5
6
7
8
9
### 进程内 FastMCP server:没有子进程,没有 socket
- word_count: Count whitespace-separated words.
- reset_counter: Pretend to wipe a counter.

### 按 destructive_hint 过滤,只留只读工具
全部 ['word_count', 'reset_counter'] -> 只读 ['word_count']

### 直接调用
word_count('a b c d') -> 4

第二种是 Path 加一个最小脚本,也就是前面 stdio 那段,验证的是真实的进程拉起和协议握手。

第三种是接上模型跑整条链。这套组合(本地 stdio server + DeepSeek + create_agent)在本机是通的:

1
2
3
4
async with MCPAdapter(SERVER) as adapter:
tools = await adapter.list_tools()
agent = create_agent(model, tools)
result = await agent.ainvoke({"messages": [{"role": "user", "content": QUESTION}]})

下面这段做了截断,只留能说明问题的几行:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
### 一次 agent 运行的完整消息序列
[0] HumanMessage
content: 先查 Oslo 的天气,再算 10 除以 0。两次工具返回的原文都要贴出来。
[1] AIMessage
tool_call: get_forecast({'city': 'Oslo'})
tool_call: divide({'a': 10, 'b': 0})
[2] ToolMessage
status: success
content: {"city":"Oslo","temp_c":3,"summary":"cloudy"}
artifact: {'structured_content': {'city': 'Oslo', 'temp_c': 3, 'summary': 'cloudy'}}
[3] ToolMessage
status: error
content: Error calling tool 'divide': division by zero
[4] AIMessage
content: 两次调用的原文返回如下:1. Oslo 天气 get_forecast:
{"city":"Oslo","temp_c":3,"summary":"cloudy"}
2. 10 除以 0 divide:Error calling tool 'divide': division by zero

模型一口气发了两个 tool_call,成功的和失败的各回一条 ToolMessage,它照着两份原文组织答案。到这一步,MCP 这条链在你的机器上就算通了,换远程 server 只是把 Path 换成 URL 加凭据。

小结

  • MCP 把工具的声明和调用协议化了。接入成本从「每个数据源写一份适配」降到「换一个 target」。
  • LangChain 这边的入口是 langchain.mcp.MCPAdapter,list_tools() 拿到的就是普通 LangChain 工具,可以直接进 create_agent。它还是 beta,记得锁版本。
  • 传输由 target 的类型决定:Path 走 stdio,http/https 字符串走 Streamable HTTP,FastMCP 实例走进程内。本地脚本一律用 Path,字符串会被当成 URL 校验。
  • 默认每个工具调用开一次会话。要跨调用复用就把 agent 调用包在 adapter 上下文里;部署时按 run 发现工具目录,底下用连接池和 cache_mode 扛开销。
  • 服务端要输入只有 elicitation 能答,落到 LangGraph interrupt。服务端要用 guard 模式返回 InputRequiredResult,且恢复会重跑工具,副作用要做幂等。