Paul's Tradecraft · Cases · AI Team · English version ↓

CASE · HOW I RUN AI AGENTS

我怎麼帶一支 AI 團隊做事

這個作品集,是我和幾個 AI agent 一起做出來的。AI 讓工作變快;但快不等於對,也不等於被授權。這頁說明我用什麼規則,讓 AI 團隊做得快,又不會替我做不該做的決定。

真實運作方法 · 非合成案例 限定範圍的操作安全 · 非全機防護
談談你公司的 AI 導入邊界 回到 How I Work

Team

Codex、DeepSeek Harness、Claude Code 為主要執行者;ChatGPT 等協助研究與規格。

Owner

我是唯一的決策者。任何角色標籤都不代表授權。

Real risk

不是 AI 有惡意,而是 AI 很有把握地做錯、用舊資訊行動,或被網頁內容誘導。

Human gate

推送、部署、對外溝通、花錢、碰密鑰、不可逆動作,一律要我明確同意。

五條規則

  1. 事實、推論、未知分開寫 — 證據不夠就標 UNKNOWN;UNKNOWN 永遠不被回報成成功。
  2. 管「效果」,不管「工具名稱」 — 同一個點擊可能只是查看,也可能是送出。依效果分四級:查看/準備/提交/控制;提交與控制需要我核准,無法判斷時視為未知並停下。
  3. 角色不帶權力 — 任何 agent 都可以實作、驗證或挑戰,但沒有一個 agent 因為角色而取得決定權。
  4. 獨立複查 — 由另一個 agent 在不讀原作者推理的情況下重做檢查,回報 PASS/PASS_WITH_GAPS/FAIL 並附證據。
  5. agent 之間的共識不是證據 — 兩個 AI 同意,只代表它們同意;是否屬實,要回到原始資料與實際測試。

一個真實例子:兩個 agent 審同一題

我請兩個 agent 各自回答同一個問題:這套系統現階段該追求哪一種安全目標?兩份審查在不互相參照的情況下得出一致結論;但其中一份找到另一份漏掉的三個證據問題,例如引用了過期的測試數字。之後由第三個 agent 把兩份並排對照,並在實機上重新確認哪些問題今天仍然存在。

哪一份成為正式版本、要不要修,不由任何 agent 決定——對照文件只把決定事項列清楚,交給我。

一個真實機制:授權閘道

對於已接上的動作,agent 不能直接執行,必須先通過一道檢查:授權要簽章、綁定到確切的任務、動作與參數,只能用一次,時效過了或內容被改過就拒絕。結果一律留下結構化收據;結果不明時記為未知,不記為成功。

Evidence ladder

可複核事實(FACT)來源未驗證/UNKNOWN
授權閘道回歸測試 30/30、轉接層測試 18/18、子行程測試 5/5、13 個強制執行情境、兩輪對抗式稽核 0 缺口(2026-09-29 本機重跑)私有治理 repo 的測試套件在測試組合以外的路徑是否同樣被攔截
兩份獨立安全審查與一份對照文件私有治理 repo外部第三方稽核
這個網站由多個 agent 建置、人工審核後部署公開 GitHub 提交紀錄這套方法在其他公司的效果

這頁不主張

不是全機或作業系統層級的防護。有 shell 或瀏覽器自動化權限的 agent,可以走不經過閘道的路徑;這套方法防的是「合作但會犯錯」的 agent,不是蓄意繞過的攻擊者。

不是沙箱、防毒或端點防護,也不宣稱能阻擋惡意程式。

治理原始碼與內部紀錄為私有;這頁只公開方法與可複核的範圍。

對公司的意義

多數公司導入 AI,第一個問題是「它能做什麼」。我先問的是:它用誰的身分、能碰哪些資料、哪些動作要人核准、出錯時怎麼知道。能力不等於授權——這是我帶進任何團隊的第一條規則。

CASE · HOW I RUN AI AGENTS

How I run a team of AI agents

This portfolio was built by me working with several AI agents. AI makes the work faster; fast is not the same as correct, and it is not the same as authorized. These are the rules that keep the team fast without letting it make decisions that are mine.

Real operating method · not a synthetic case Bounded operational safety · not host-wide protection

Team

Codex, DeepSeek Harness, and Claude Code are the primary runtimes; ChatGPT and others help with research and specs.

Owner

I am the only decision-maker. No role label confers authority.

Real risk

Not malicious AI, but AI that is confidently wrong, acts on stale information, or is steered by web content.

Human gate

Push, deploy, external communication, spend, secrets, and irreversible actions always need my explicit approval.

Five rules

  1. Facts, inferences, and unknowns stay separate — without evidence it is UNKNOWN, and UNKNOWN is never reported as success.
  2. Govern effects, not tool names — the same click can be a look or a submission. Four effect classes: read / prepare / commit / control. Commit and control need my approval; anything undeterminable is treated as unknown and stops.
  3. A role carries no authority — any agent may build, verify, or challenge; none gains decision rights from its role.
  4. Independent re-checks — another agent redoes the check without reading the producer's reasoning and reports PASS / PASS_WITH_GAPS / FAIL with evidence.
  5. Agreement between agents is not evidence — two AIs agreeing only means they agree; truth comes from original sources and real tests.

A real example: two agents, one question

Two agents answered the same question independently: what security objective should this system pursue now? They reached the same conclusion without consulting each other, but one found three evidence problems the other missed, such as citing test counts from an older run. A third agent then placed both side by side and re-checked on the actual machine which problems still held. Which review becomes canonical, and what gets fixed, is not an agent's call: the reconciliation lists the decisions and hands them to me.

A real mechanism: the authorization gateway

For connected actions, an agent cannot execute directly. Authorization must be signed, bound to the exact task, action, and arguments, used once, and rejected if expired or altered. Every result leaves a structured receipt; an unclear result is recorded as unknown, never as success.

Evidence ladder

Verifiable factSourceNot verified / UNKNOWN
Authorization gateway regression tests 30/30, adapter tests 18/18, subprocess tests 5/5, 13 enforcement scenarios, two adversarial audit rounds with 0 gaps (re-run locally 2026-09-29)Test suites in the private governance repoWhether paths outside the tested set are intercepted the same way
Two independent security reviews and one reconciliation documentPrivate governance repoExternal third-party audit
This site was built by several agents and deployed after human reviewPublic GitHub commit historyHow well the method works at other companies

What this page does not claim

It is not host-wide or OS-level protection. An agent with shell or browser automation can take paths that never pass the gateway; the method protects against cooperative-but-fallible agents, not a deliberate attacker.

It is not a sandbox, antivirus, or endpoint protection, and makes no malware-blocking claim.

The governance source and internal records are private; this page publishes the method and its verifiable scope only.

Why it matters to a company

Most AI adoption starts with "what can it do?" I start with: whose identity does it act under, what data may it touch, which actions need a human, and how would we know when it is wrong. Capability is not authority — that is the first rule I bring into any team.