传统代码补全的输入是当前文件的光标上下文,输出是一段建议代码。编码智能体(Codex 类)的输入是整个任务,输出是一组跨文件的变更,并且它需要主动收集信息、执行命令、验证结果。
这个差异决定了工程上要解决四类新问题:
问题 | 说明 |
|---|---|
上下文 | 仓库很大,不能全塞进模型 |
执行 | 智能体要能跑命令、跑测试 |
审查 | 产出是多文件 diff,必须可读 |
边界 | 不能让它在生产环境乱执行 |
本文从这四个问题出发,给出可运行骨架。
仓库越大,上下文选择越关键。把整个仓库塞进模型既贵又噪,反而降低准确率。正确做法是先按关键词打分,再按 token 预算裁剪。
from dataclasses import dataclass
from pathlib import Path
IGNORE_DIRS = {".git", "node_modules", "__pycache__", ".venv",
"dist", "build", ".next"}
ALLOW_EXT = {".py", ".ts", ".tsx", ".go", ".java",
".md", ".toml", ".yaml", ".yml"}
@dataclass
class FileMeta:
path: str
tokens: int
def scan_repo(root: Path, max_files: int = 3000) -> list[FileMeta]:
files = []
for p in root.rglob("*"):
if not p.is_file() or p.suffix not in ALLOW_EXT:
continue
if any(part in IGNORE_DIRS for part in p.parts):
continue
text = p.read_text(encoding="utf-8", errors="ignore")
files.append(FileMeta(
str(p.relative_to(root)), len(text) // 4))
if len(files) >= max_files:
break
return files
def select_context(files: list[FileMeta], keywords: list[str],
budget: int = 60000) -> list[str]:
scored = []
for f in files:
score = sum(1 for k in keywords if k.lower() in f.path.lower())
scored.append((score, f))
scored.sort(key=lambda x: (-x[0], x[1].tokens))
chosen, used = [], 0
for _, f in scored:
if used + f.tokens > budget:
continue
chosen.append(f.path)
used += f.tokens
return chosen两个设计点值得说明:
智能体要能跑测试、跑 lint,但不能让它执行任意命令。沙箱的核心是命令白名单 + 工作目录隔离 + 超时控制。
import subprocess
from pathlib import Path
ALLOWED_CMDS = {"ruff", "mypy", "pytest", "npm", "go", "cargo"}
FORBIDDEN_ARGS = ("rm", "-rf", "sudo", "curl", "wget", "chmod")
def safe_run(cmd: list[str], cwd: Path,
timeout: int = 120) -> dict:
if not cmd or cmd[0] not in ALLOWED_CMDS:
return {"code": -1, "err": f"命令不在白名单:{cmd[:1]}"}
joined = " ".join(cmd).lower()
if any(f in joined for f in FORBIDDEN_ARGS):
return {"code": -1, "err": "命中禁止参数"}
try:
p = subprocess.run(
cmd, cwd=cwd, capture_output=True, text=True,
timeout=timeout, shell=False,
)
return {"code": p.returncode,
"out": p.stdout[-4000:],
"err": p.stderr[-2000:]}
except subprocess.TimeoutExpired:
return {"code": -1, "err": "timeout"}三个关键约束:
shell=False:避免 shell 注入。智能体产出的是多文件变更,直接合并风险极高。必须先转成 diff,再做结构化审查。
import difflib
from pathlib import Path
def make_diff(old: str, new: str, path: str) -> str:
return "".join(difflib.unified_diff(
old.splitlines(keepends=True),
new.splitlines(keepends=True),
fromfile=f"a/{path}", tofile=f"b/{path}",
))
def summarize(diff: str) -> dict:
added = sum(1 for l in diff.splitlines()
if l.startswith("+") and not l.startswith("+++"))
removed = sum(1 for l in diff.splitlines()
if l.startswith("-") and not l.startswith("---"))
files = diff.count("--- ")
return {"added": added, "removed": removed, "files": files}
RISKY_PATTERNS = [
("eval(", "疑似动态执行"),
("shell=True", "疑似 shell 注入"),
("verify=False", "疑似关闭证书校验"),
("# type: ignore", "疑似绕过类型检查"),
("except: pass", "疑似吞异常"),
]
def risk_scan(diff: str) -> list[str]:
flags = []
for line in diff.splitlines():
if not line.startswith("+"):
continue
for pattern, desc in RISKY_PATTERNS:
if pattern in line:
flags.append(f"{desc}: {line.strip()[:60]}")
return flags审查清单要覆盖:是否改了无关文件、是否引入新依赖、是否吞异常、是否绕过类型检查、是否硬编码密钥、是否删除测试。
name: ci
on: [push, pull_request]
jobs:
quality:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install ruff mypy pytest bandit pip-audit
- run: ruff check .
- run: mypy .
- run: pytest -q
- run: bandit -r app
- run: pip-auditlint、类型检查、测试、安全扫描、依赖审计,一个都不能少。任何"这是 AI 生成的所以免检"的想法,都是把风险转嫁给生产环境。
import time, logging
log = logging.getLogger("codex-agent")
def agent_loop(plan: list[dict], repo: Path,
apply_patch, max_rounds: int = 5) -> list[dict]:
for rnd in range(max_rounds):
pending = [s for s in plan if s["status"] != "done"]
if not pending:
break
for step in pending:
start = time.time()
patch = apply_patch(step) # 模型生成补丁
checks = {
"lint": safe_run(["ruff", "check", "."], repo),
"type": safe_run(["mypy", "."], repo),
"test": safe_run(["pytest", "-q"], repo),
}
ok = all(c["code"] == 0 for c in checks.values())
if ok:
step["status"] = "done"
else:
step["status"] = "retry"
failed = [k for k, v in checks.items() if v["code"] != 0]
step["goal"] += f"\n上次失败:{failed}"
log.info("round=%d step=%s ok=%s cost=%.2fs",
rnd, step.get("id"), ok, time.time() - start)
return plan循环的核心是失败信息回填。测试失败时,把失败信息追加到下一步提示词,让智能体自我修正。轮次要有上限,避免无限循环烧钱。
维度 | 做法 | 缺失后果 |
|---|---|---|
上下文 | 关键词打分 + token 预算 | 噪声大、成本高 |
执行 | 命令白名单 + shell=False | 注入、误删 |
隔离 | 仓库副本 + 超时 | 污染生产环境 |
审查 | diff + 风险扫描 | 危险变更被合并 |
门禁 | lint/mypy/pytest/bandit | 质量问题流入主干 |
观测 | trace、轮次、耗时、工具调用 | 问题无法定位 |
合规底线:不把密钥、用户数据、内部源码提交给外部模型;智能体只在沙箱和独立分支执行;所有合并需人工确认;AI 生成代码按团队规范标注。
Codex 智能体的工程核心,是把智能体放进一条可控的工程链路:
代码可以简单,但白名单、隔离、审查、门禁这四件事不能省。智能体能放大产能,但架构判断、安全责任和合规义务仍然在人。
原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。