代码仓库ChainReaction

写了几个 agent 之后,让人卡住的通常不是工具调用怎么写。麻烦的是”这次运行该拿到什么数据”:用户 ID 从哪来,上一轮对话的偏好存在哪,工具返回的 8000 字文档要不要整个塞回模型。这些问题在 create_agent 的入门示例里看不到,因为示例里只有一个 messages。

LangChain 把答案拆成了三块:state、context、store。名字都短,含义却经常被混着用。我见过把 user_id 塞进 state 的写法,也见过每次 invoke 都把数据库连接写进 messages 的写法。

这篇先把这三块拆开,再看 context_schema 和 runtime.context 怎么用,然后把 runtime 上剩下的字段挨个跑一遍。剩下的一节讲上下文工程:窗口怎么被填满,又该怎么打扫。

文中的输出都是在 Python 3.12 上跑 DeepSeek 得到的真实终端结果,脚本放在 Runtime/。

三个数据源

官方文档在 context engineering 页面给了一张表,按作用域把三者分开了。我补上读写方式:

数据源 文档里的别称 作用域 生命周期 典型内容
Runtime Context 静态配置 单次运行 invoke 传入,运行结束即失效 user_id、API key、数据库连接、权限
State 短期记忆 一个线程 由 checkpointer 持久化,跨轮次累积 messages、上传的文件、认证状态、工具结果
Store 长期记忆 跨线程 存储还在就一直有 用户偏好、抽取出的记忆、历史数据

Runtime Context 是只读的。一次 invoke 传进去,整个运行里只能读,工具和中间件都改不了它。State 可变,工具返回 Command(update=...) 或者中间件返回一个 dict,都会被图的 reducer 合并回去。Store 要显式 get/put,写进去的东西不会自动出现在模型的消息列表里。

三者在一次运行里的流向大致是这样:

工具和中间件都能读这三块,能写回去的只有 State 和 Store。

还有一对容易混的:thread_id 和 context。前者通过 config 传,管的是对话本身,也就是消息历史和 checkpoint 归哪个线程;后者通过 context 传,管的是这次运行要用的数据。文档在 agents 页面里说得很直白,thread_id 划定对话的范围,context 承载单次运行的数据,两者经常一起传。

我写了个脚本把三块同时打印出来。为了让 State 的”线程内”这个属性可见,agent 加了 checkpointer,invoke 时带上 thread_id:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
# Runtime/RuntimeFields.py(节选)
import os
from langchain_openai import ChatOpenAI

# 全篇共用的模型配置,后面片段里出现的 model 都指这个
model = ChatOpenAI(
api_key=os.getenv('DEEPSEEK_API_KEY'),
base_url="https://api.deepseek.com/v1",
model="deepseek-chat",
temperature=0.1,
max_tokens=1000,
)

agent = create_agent(
model=model,
tools=[inspect_tool_runtime],
middleware=[dump_runtime_fields],
context_schema=UserContext,
store=store,
checkpointer=InMemorySaver(),
)

result = agent.invoke(
{"messages": [{"role": "user", "content": "看看 runtime 上有哪些字段?"}]},
config={"configurable": {"thread_id": "thread-demo-1"}},
context=UserContext(user_id="u_1024"),
)

真实输出,省略了对象地址:

1
2
3
4
5
6
7
8
9
10
11
12
=== [tool] ToolRuntime 实例上每个字段的真实取值 ===
runtime.state = {'messages': [HumanMessage(content='看看 runtime 上有哪些字段?', ...
runtime.context = UserContext(user_id='u_1024')
runtime.config = {'tags': [], 'metadata': {'ls_integration': 'langchain_create_agent', 'thread_id': 'thread...
runtime.stream_writer = <function Pregel.stream.<locals>.stream_writer at 0x...>
runtime.tool_call_id = 'call_00_4NfVsUPhIOXRSqRf06b69410'
runtime.store = <langgraph.store.memory.InMemoryStore object at 0x...>
runtime.tools = [StructuredTool(name='inspect_tool_runtime', description='打印工具里能看到的 runtime 字段。', ...
runtime.execution_info = ExecutionInfo(checkpoint_id='1f1c0080-1c42-6329-8002-cdb734816007', ...
runtime.server_info = None
runtime.tools 里的工具名 = ['inspect_tool_runtime']
runtime.store.get(('prefs',), 'u_1024') = Item(namespace=['prefs'], key='u_1024', value={'tone': 'concise'}, created_at='...', updated_at='...')

一次工具调用里三块同时在线,来源各不相同:context 来自 invoke 的参数,state 来自这个线程的 checkpoint,store 来自你传给 create_agent 的那个存储实例。

context_schema 与 runtime.context

context_schema 是一个类型声明,告诉 create_agent 这次运行的 context 长什么样。文档用 dataclass 举例:

1
2
3
4
5
6
7
from dataclasses import dataclass

@dataclass
class UserContext:
user_id: str
plan: str
locale: str

在工具里读它,靠 ToolRuntime 参数。参数名必须叫 runtime,这个名字和 config 都是保留的,自己定义工具参数时不能占用,否则运行时报错。

1
2
3
4
5
6
7
8
9
10
# Runtime/ContextTool.py
@tool
def get_account_profile(runtime: ToolRuntime[UserContext]) -> str:
"""查询当前用户的账户套餐信息。"""
ctx = runtime.context
log(f"[tool] type(runtime.context) = {type(ctx).__name__}")
log(f"[tool] user_id = {ctx.user_id}")
log(f"[tool] plan = {ctx.plan}")
log(f"[tool] locale = {ctx.locale}")
return f"user_id={ctx.user_id}; plan={ctx.plan}; locale={ctx.locale}"

注意这个工具没有 user_id 参数。模型看不到它,也没有机会去编造它。真实运行输出:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
[tool] type(runtime) = ToolRuntime
[tool] type(runtime.context) = UserContext
[tool] user_id = u_1024
[tool] plan = pro
[tool] locale = zh-CN

=== 消息序列 ===
HumanMessage: 帮我查一下我现在是什么套餐?
AIMessage: tool_calls=['get_account_profile']
ToolMessage: user_id=u_1024; plan=pro; locale=zh-CN
AIMessage: 您当前的套餐是 Pro 版。

=== 最终回答 ===
您当前的套餐是 Pro 版。

中间件里读 context 有两个入口。节点式钩子(before_model、after_model)直接收 Runtime 参数,包裹式钩子(wrap_model_call)从 request.runtime 拿。文档给的动态提示写法是 @dynamic_prompt:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
# Runtime/DynamicSystemPrompt.py
@dynamic_prompt
def context_aware_prompt(request: ModelRequest) -> str:
ctx = request.runtime.context
base = "你是一个中文技术助手,回答控制在两句话以内。"

if ctx.role == "admin":
base += "\n当前用户是管理员,可以给出内部命令和完整配置项。"
else:
base += "\n当前用户是普通读者,只讲公开用法,不要提内部配置。"

if ctx.deployment_env == "production":
base += "\n当前是生产环境,涉及写操作要先提醒风险。"

base += f"\n请称呼用户为 {ctx.user_name}。"
return base

同一个 agent,同一句提问,换掉 context 之后系统提示和回答都跟着变了:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
=== 运行:role=admin, env=production ===
[middleware] role=admin env=production -> 注入的 system prompt:
你是一个中文技术助手,回答控制在两句话以内。
当前用户是管理员,可以给出内部命令和完整配置项。
当前是生产环境,涉及写操作要先提醒风险。
请称呼用户为 老张。
回答: 老张,生产环境重启前请先确认:修改 service.conf 中的 restart_timeout 和 graceful_shutdown 参数,并备份当前配置。重启有服务中断风险,建议在低峰期操作并提前通知相关方。

=== 运行:role=viewer, env=dev ===
[middleware] role=viewer env=dev -> 注入的 system prompt:
你是一个中文技术助手,回答控制在两句话以内。
当前用户是普通读者,只讲公开用法,不要提内部配置。
请称呼用户为 老张。
回答: 老张,重启线上服务一般只需在控制台点"重启"按钮,或执行服务自带的 restart 命令即可。具体配置项请以你所用平台的官方文档为准。

有个坑我踩过。context_schema 只做声明,不传 context 不会在 create_agent 阶段报错,也不会在 invoke 入口报错,要等你真去读字段的时候才炸:

1
2
=== 忘记传 context 会怎样 ===
AttributeError: 'NoneType' object has no attribute 'role'

报错点在中间件里那一行 ctx.role。如果这段逻辑藏在几层调用之外的工具里,你会先看到一条工具执行失败的记录,再往回找。稳妥的做法是在 before_agent 里加一句校验,context 缺字段就直接抛 ValueError,别让它跑进模型循环。

为什么静态依赖不该塞进 state

user_id、数据库连接、API key 这些东西塞进 state 也能跑,但要付三个代价。

State 会被 checkpointer 序列化。文档写得很清楚,state 通过 checkpointer 持久化到数据库或内存,这样线程随时能恢复。把数据库连接对象放进去,等于要求它可序列化,还得跟着线程活多久就活多久。连接池不这么用。

State 的作用域是线程,配置的作用域通常是整个应用。同一份 deployment_env 服务所有线程,放进 state 就得在每次开新线程时重复写一遍,每个线程各存一份副本,改配置时找不到单一来源。

还有个更隐蔽的问题。State 里的 messages 会直接进模型上下文,其他自定义字段默认不进 prompt,但只要有人写了个把 state 整体注入的中间件,敏感配置就跟着出去了。context 是显式的按运行传入,评审代码时能一眼看到谁传了什么。测试的时候也可以传一个假的 context 对象,不用真的连数据库。

文档在 runtime 页面用一句话概括了这件事:runtime context 提供的是依赖注入。不用硬编码,也不用全局变量,需要什么就在 invoke 时传进来。这样同一个 agent 在测试里拿到假连接,在生产里拿到真连接,中间一行代码都不用改。

runtime 上还有什么

文档在 tools 页面列了 ToolRuntime 能拿到的东西,一共八项:

组件 说明 典型用途
State 当前会话的可变数据 读对话历史,累计工具调用次数
Context invoke 时传入的只读配置 按用户身份调整回答
Store 跨会话的持久数据 存用户偏好,维护知识库
Stream Writer 工具执行期间发实时更新 长任务进度
Execution Info 本次执行的标识与重试信息 thread_id、run_id、第几次尝试
Server Info LangGraph Server 的元数据 assistant ID、认证用户
Config 本次执行的 RunnableConfig 回调、tags、metadata
Tool Call ID 本次工具调用的唯一标识 关联日志,拼 ToolMessage

这是”能用来做什么”的清单,实际字段名要跑一下才知道。我把两个 dataclass 的字段都打印了:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
=== ToolRuntime 的字段(langchain.tools.ToolRuntime)===
state StateT
context ContextT
config RunnableConfig
stream_writer StreamWriter
tool_call_id str | None
store BaseStore | None
tools list[BaseTool]
execution_info ExecutionInfo | None
server_info ServerInfo | None

=== Runtime 的字段(langgraph.runtime.Runtime)===
context ContextT
store BaseStore | None
stream_writer StreamWriter
heartbeat Callable[[], None]
previous Any
execution_info ExecutionInfo | None
server_info ServerInfo | None
control RunControl | None

=== ExecutionInfo 的字段 ===
checkpoint_id str
checkpoint_ns str
task_id str
thread_id str | None
run_id str | None
node_attempt int
node_first_attempt_time float | None

ToolRuntime 比 Runtime 多出 state、config、tool_call_id、tools。这符合直觉:中间件本来就在图里,拿得到 state;工具不在,得从 runtime 里补。tools 这个字段文档没提,实际是当前注册的工具列表,写权限过滤逻辑时可能用得上。

config 文档归在 ToolRuntime 那一栏,实际两个 runtime 上都有。它装的是本次执行的 RunnableConfig,从我的输出里能看到 metadata 里带着 ls_integration 和 thread_id,回调、tags 也在里面。要在工具里触发自定义回调或者按 tag 分流时才会用到它。

反过来,Runtime 上有三个字段文档没写:heartbeat、previous、control。我在本地 langgraph 1.2.12 上确认它们存在,但文档既然没承诺,我不建议依赖。heartbeat 打印出来是 _no_op_heartbeat,本地运行时是个空实现。

stream_writer

工具跑长任务时,让用户盯着一个转圈图标很难受。runtime.stream_writer 让工具往 custom 流里推消息:

1
2
3
4
5
6
7
8
9
# Runtime/StreamWriter.py
@tool
def index_documents(topic: str, runtime: ToolRuntime[UserContext]) -> str:
"""为指定主题建立索引,过程中通过 stream_writer 汇报进度。"""
writer = runtime.stream_writer
total = 3
for i in range(1, total + 1):
writer({"progress": f"{i}/{total}", "stage": f"embedding {topic} chunk {i}"})
return f"{topic} 的 {total} 个分片已建好索引"

消费端在 stream_mode 里加上 "custom":

1
2
3
4
5
6
7
8
9
10
11
12
for chunk in agent.stream(
{"messages": [{"role": "user", "content": "帮我给 langchain-runtime 建个索引"}]},
stream_mode=["custom", "updates"],
version="v2",
context=UserContext(user_id="u_1024"),
):
if chunk["type"] == "custom":
log(f"custom <- {chunk['data']}")
elif chunk["type"] == "updates":
for node, update in chunk["data"].items():
for msg in update.get("messages", []):
...

真实输出:

1
2
3
4
5
6
7
=== agent.stream(stream_mode=['custom', 'updates']) 收到的分片 ===
updates <- [model] I'll create an index for langchain-runtime.
custom <- {'progress': '1/3', 'stage': 'embedding langchain-runtime chunk 1'}
custom <- {'progress': '2/3', 'stage': 'embedding langchain-runtime chunk 2'}
custom <- {'progress': '3/3', 'stage': 'embedding langchain-runtime chunk 3'}
updates <- [tools] langchain-runtime 的 3 个分片已建好索引
updates <- [model] 已为 langchain-runtime 建好索引,共 3 个分片。

custom 分片和 updates 分片交错到达,顺序就是实际执行顺序。

server_info:本地是 None,起个 server 才有值

文档在 tools 页面写得很短:工具跑在 LangGraph Server 上时,用 runtime.server_info 拿 assistant ID、graph ID 和认证用户,本地开发时它是 None。前面 runtime_fields.py 输出里那行 runtime.server_info = None 就是这个情况。我想知道它真有值时长什么样,于是在本地起了一个 langgraph dev。

工程三个文件,放在 Runtime/langgraph_dev/。langgraph.json 声明图入口:

1
2
3
4
5
6
7
{
"dependencies": ["."],
"graphs": {
"server_info_agent": "./agent.py:agent"
},
"python_version": "3.12"
}

agent.py 里就是个普通的 create_agent,中间件和工具各打印一次 server_info,模型还是前面那套 DeepSeek 配置:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
# Runtime/langgraph_dev/agent.py(节选)
import dataclasses

from langchain.agents import AgentState, create_agent
from langchain.agents.middleware import before_model
from langchain.tools import ToolRuntime, tool
from langgraph.runtime import Runtime


def dump(where: str, runtime) -> None:
"""打印 server_info 本体和它的每个字段。"""
info = runtime.server_info
log(f"[{where}] runtime.server_info = {info!r}")
if info is None:
log(f"[{where}] server_info 是 None,没有字段可打印")
return
for f in dataclasses.fields(info):
log(f"[{where}] server_info.{f.name} = {getattr(info, f.name)!r}")


@before_model
def probe_before_model(state: AgentState, runtime: Runtime) -> dict | None:
dump("before_model", runtime)
return None


@tool
def report_server_info(runtime: ToolRuntime) -> str:
"""报告当前这次运行所在 LangGraph Server 的元数据。"""
dump("tool", runtime)
info = runtime.server_info
return f"assistant_id={info.assistant_id}; graph_id={info.graph_id}; user={info.user}"

起服务就一条命令,默认监听 127.0.0.1:2024:

1
langgraph dev --no-browser --no-reload --port 2024

客户端用 langgraph_sdk 打过去。这个脚本会把同一个 agent 先在本地直接 invoke 一次,再通过 server 跑一次:

1
2
3
4
5
6
7
8
9
10
# Runtime/ServerInfo.py(节选)
client = get_client(url="http://127.0.0.1:2024")
thread = await client.threads.create()
async for chunk in client.runs.stream(
thread["thread_id"],
"server_info_agent",
input={"messages": [{"role": "user", "content": "这次运行跑在哪个 server 上?"}]},
stream_mode="values",
):
...

两边的探针输出放在一起,差别很直接:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
=== 1. 本地直接 invoke 同一个 agent ===
=== [middleware/before_model] ===
[before_model] runtime.server_info = None
[before_model] server_info 是 None,没有字段可打印

=== [tool] ===
[tool] runtime.server_info = None
[tool] server_info 是 None,没有字段可打印

=== 2. 连上 http://127.0.0.1:2024 ===
assistant_id=362004ad-4adb-5e16-bc77-1fe6e22b2661 graph_id=server_info_agent name=server_info_agent
thread_id = 01a107be-c9f2-76a1-9e8c-7059ed8185f3

=== 3. 这次 run 的消息序列(客户端看到的)===
human: 这次运行跑在哪个 server 上?
ai:
tool_call -> report_server_info({})
tool: assistant_id=362004ad-4adb-5e16-bc77-1fe6e22b2661; graph_id=server_info_agent; user=None
ai: 这次运行所在的 server 信息为:assistant_id=362004ad-4adb-5e16-bc77-1fe6e22b2661,graph_id=server_info_agent,user 为空。

=== 4. 服务端进程里打印的 runtime.server_info ===
=== [middleware/before_model] ===
[before_model] runtime.server_info = ServerInfo(assistant_id='362004ad-4adb-5e16-bc77-1fe6e22b2661', graph_id='server_info_agent', user=None)
[before_model] server_info.assistant_id = '362004ad-4adb-5e16-bc77-1fe6e22b2661'
[before_model] server_info.graph_id = 'server_info_agent'
[before_model] server_info.user = None

=== [tool] ===
(工具里打印的内容和中间件一样,省略)

ServerInfo 就三个字段。assistant_id 是图在这个 server 上的注册 ID,我起了三次服务,它每次都是同一个值;graph_id 就是 langgraph.json 里 graphs 那个键。user 是认证用户,文档里的用法是 server.user.identity,我这边没配 auth,所以是 None。要让它有值,得给 langgraph.json 配 auth,或者用平台的认证。

两个细节记一下。

中间件和工具里看到的 server_info 是同一份值,assistant_id 完全一致。它描述的是”这次运行落在哪个 server 上”,不是某个节点的局部状态。

langgraph dev 在中文 Windows 上第一次没起来。langgraph-api 0.15.1 的 validation.py 第 14 行是 open(pathlib.Path(__file__).parent.parent / "openapi.json"),没指定编码,GBK 环境下按 GBK 去解这个 UTF-8 文件,直接崩在 import 阶段:

1
2
3
File "...\site-packages\langgraph_api\validation.py", line 14, in <module>
openapi_str = f.read()
UnicodeDecodeError: 'gbk' codec can't decode byte 0x94 in position 1660: illegal multibyte sequence

加一句 $env:PYTHONUTF8='1' 再起就正常了。这跟 agent 里写什么没关系,是服务端自己的编码问题。

几个坑

execution_info 需要 langgraph>=1.1.5,文档在 runtime 页面标注了这一点。我这边装的是 1.2.12。

thread_id 要有 checkpointer 才有值。第一次跑这个脚本时我没加 checkpointer,打印出来是 runtime.execution_info.thread_id = None;加上 InMemorySaver 并传 thread_id 之后才变成 thread-demo-1。

工具里用 runtime.tool_call_id 拼 ToolMessage 时,id 必须和触发它的那次工具调用一致。文档明确写了 ToolNode 会检查,缺了会抛 ValueError。自己构造 ToolMessage 时记得把这个字段带上。

runtime.stream_writer 必须在 LangGraph 的执行上下文里调用,脱离图直接 tool.invoke() 会失败。

上下文工程的实操

前面是”能拿到什么”,这一段是”拿到之后怎么放”。文档把可控的东西分成模型上下文(instructions、messages、tools、model、response format)和生命周期上下文,前者每次调用临时生效,后者会落到 state 里。这个区分决定了你该选哪个钩子。

system prompt 分层

把系统提示拆成固定骨架和动态片段。骨架放 create_agent 的 system_prompt,动态片段在 @dynamic_prompt 里拼,上一节的 role/env 例子就是这个模式。文档里还有一种写法是在 wrap_model_call 里改 request.system_message,用 content_blocks 追加而不是覆盖:

1
2
3
4
5
6
7
8
from langchain.messages import SystemMessage

@wrap_model_call
def add_context(request: ModelRequest, handler) -> ModelResponse:
new_content = list(request.system_message.content_blocks) + [
{"type": "text", "text": "Additional context."}
]
return handler(request.override(system_message=SystemMessage(content=new_content)))

文档特别提示 request.system_message 永远是 SystemMessage 对象,哪怕你创建 agent 时传的是字符串。想追加内容就走 content_blocks,直接拼字符串会把原有结构弄丢。

消息裁剪

裁剪是临时的:只改这一次发给模型的消息,state 不动。用 wrap_model_call 加 request.override(messages=...):

1
2
3
4
5
6
7
8
9
@wrap_model_call
def trim_messages(request: ModelRequest, handler) -> ModelResponse:
messages = list(request.messages)
if len(messages) > 4:
kept = messages[-4:]
while kept and not isinstance(kept[0], HumanMessage):
kept = kept[1:]
request = request.override(messages=kept)
return handler(request)

那几行 while 不是凑数的。消息列表里有 AIMessage(tool_calls=...) 和对应的 ToolMessage,从中间切会把它们拆散,模型接口会直接报错。切完要让第一条是 HumanMessage。

按需注入

有些上下文只在特定条件下才需要。文档给的两个例子是按用户上传的文件摘要注入、按用户所在辖区注入合规条款,两个都用 wrap_model_call 把新消息拼到列表末尾。文档里有句话值得记住:模型对末尾的消息更敏感,所以注入的内容放最后,别塞进开头。

注入用 request.override(messages=...) 拼,state 不受影响。如果注入的信息需要被后续轮次记住,那就得换节点式钩子,返回一个 dict 让它落进 state。

工具结果压缩

工具返回 8000 字文档,模型只需要其中一段。压缩放在 wrap_tool_call 里,拿到结果后重建一个 ToolMessage:

1
2
3
4
5
6
7
8
9
10
11
@wrap_tool_call
def compress_tool_result(request, handler) -> ToolMessage | Command:
result = handler(request)
if isinstance(result, ToolMessage) and isinstance(result.content, str):
original = len(result.content)
if original > 200:
result = ToolMessage(
content=result.content[:200] + f"...(已压缩,原始 {original} 字符)",
tool_call_id=result.tool_call_id,
)
return result

原地改 result.content 其实也能跑通,ToolMessage 的 model_config 是 {'extra': 'allow'},不是冻结模型。我选择重建,是为了把 tool_call_id 这类字段显式带上,压缩逻辑变复杂时不容易漏。

跑一遍看效果

三个技巧放进同一个 agent,跑三轮对话:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
=== 第 1 轮:你好,先打个招呼。 ===
[wrap_model_call] 消息 1->1 条,字符 9->9(state 不变)
state 里的消息总数 = 2
回答: 你好!有什么可以帮你的?

=== 第 2 轮:再确认一下,你在吗? ===
[wrap_model_call] 消息 3->3 条,字符 31->31(state 不变)
state 里的消息总数 = 4
回答: 在的,随时可以提问。

=== 第 3 轮:用 search_docs 查一下 runtime context,然后一句话总结。 ===
[wrap_model_call] 消息 5->3 条,字符 83->62(state 不变)
[tool] 检索 'runtime context',原始返回 4800 字符
[wrap_tool_call] search_docs: 4800 -> 219 字符
[wrap_model_call] 消息 7->3 条,字符 302->261(state 不变)
state 里的消息总数 = 8
回答: 检索结果只返回了重复的占位片段("LangChain Runtime 文档片段"),没有实际内容,因此无法给出有依据的总结。建议换更具体的关键词...

第 3 轮发生了两次模型调用。第一次裁剪前 5 条消息,裁到 3 条;工具返回 4800 字符,压到 219 字符;第二次模型调用前 state 里已经有 7 条消息,同样裁到 3 条。整个过程 state 从 2 涨到 8,说明裁剪和压缩都只作用于这一次调用,历史一条没少。

窗口的填充和打扫大致是这个节奏:

如果目标是永久替换历史,裁剪就不合适。文档给的是 SummarizationMiddleware,超过阈值时它用另一个模型总结旧消息,把总结写回 state,后续轮次看到的是总结。我用 trigger=("messages", 3) 和 keep=("messages", 1) 跑了一遍,第四轮它还记得第一轮报的名字:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
--- 第 1 轮:我叫老张。
state 里消息 2 条
HumanMessage: 我叫老张。
AIMessage: 你好,老张!
--- 第 2 轮:我在做 LangChain 的笔记。
state 里消息 3 条
HumanMessage: Here is a summary of the conversation to date:

## SESSION INTENT
The user introduced themselves as "老张" (Lao Zhang). No specific task or goal has been stated yet.

## SUMMARY
The user provided their name: 老张 (Lao Zhang). The assistant greeted them. ...
HumanMessage: 我在做 LangChain 的笔记。
AIMessage: 好的,老张,需要我帮你整理或解答 LangChain 的哪部分内容?
--- 第 4 轮:我刚才说我叫什么?
state 里消息 3 条
HumanMessage: Here is a summary of the conversation to date:
## SUMMARY
- 用户自我介绍为"老张"(Lao Zhang)。
- 用户表示正在做 LangChain 的笔记。
- 用户补充说明自己住在杭州。
HumanMessage: 我刚才说我叫什么?
AIMessage: 你刚才说你叫"老张"。

两条原始消息(我叫老张。 和对应的回答)在第 2 轮就被换成了一个 HumanMessage 里的总结,state 长度从此稳定在 3 条。裁剪是每次调用算一遍,便宜但每次都要重算;总结要多花一次模型调用,换来的是历史真的变短了。对话只有十几轮用裁剪,涨到几百条再上总结。

什么时候用中间件,什么时候写死

情况 建议
内容对所有用户、所有运行都一样 直接写进 system_prompt
内容依赖 user_id、环境、权限 context_schema 加 @dynamic_prompt
内容依赖对话长度或历史 中间件读 request.messages
内容来自外部存储 中间件读 runtime.store,或让工具按需取
只有少数轮次需要 让工具按需返回,别提前注入
每次调用都要变,历史不能改 wrap_model_call 加 request.override()
历史本身要永久变短 SummarizationMiddleware 或 before_model 返回 dict

文档 best practices 的第一条是 Start simple:先用静态 prompt 和固定工具集,真出现了第二种变体再抽中间件。我按这个原则试过,只有两三个分支的动态 prompt 直接写在 @dynamic_prompt 里比抽成类更好读。中间件真正的成本是后来的人要跳三个文件才能拼出完整的系统提示。

另一个判断标准是数据从哪来。数据已经在 context 或 store 里,中间件就是顺手的事;数据要靠一次网络请求才能拿到,那它应该是工具,让模型自己决定什么时候花这个钱。

小结

  • Runtime.context 是单次运行的只读依赖,用 context_schema 声明类型,invoke(context=...) 传入;工具里通过 ToolRuntime 参数读,中间件里通过 Runtime 参数或 request.runtime 读。不传 context 不会提前报错,会在读字段时抛 AttributeError。
  • State 是线程内的可变数据,会被 checkpointer 持久化;Store 是跨线程的持久数据,要显式读写。数据库连接、API key 这类静态依赖放 context,不要放 state。
  • ToolRuntime 有九个字段,文档列了其中八项的用途;Runtime 上还有 heartbeat、previous、control 三个没写进文档,不建议依赖。server_info 本地直接 invoke 是 None,我起了个 langgraph dev 才拿到真实值,三个字段是 assistant_id、graph_id、user;thread_id 要有 checkpointer 才有值。
  • 上下文工程的四个常用动作:system prompt 分层、消息裁剪、按需注入、工具结果压缩。前三个用 wrap_model_call 加 request.override() 做临时修改,压缩用 wrap_tool_call 重建 ToolMessage。
  • 临时修改和持久修改是两条路。要历史真的变短就用 SummarizationMiddleware 或 before_model 返回 dict;只想让这一次调用省点 token,就用 override。

脚本都在仓库的 Runtime/ 下:ContextTool.py、DynamicSystemPrompt.py、RuntimeFields.py、StreamWriter.py、ContextWindow.py、Summarization.py。server_info 那节的服务端工程在 Runtime/langgraph_dev/,客户端脚本是 Runtime/ServerInfo.py。