Checkpoint และ Resume งาน Agent อย่างปลอดภัย
ออกแบบ “Checkpoint และ Resume งาน Agent อย่างปลอดภัย” ให้ตรวจสอบ ควบคุมต้นทุน และกู้คืนได้ในงานจริง
เป้าหมาย: เลือกและควบคุมสถาปัตยกรรมหลาย Agent จาก dependency และความเสี่ยง · ใช้เวลาประมาณ 16 นาที
Decompose → Contract → Coordinate → Recover
แยกงานเฉพาะเมื่อการแยกทำให้ตรวจและกู้คืนง่ายขึ้น
Decompose
แยกตาม dependency, data boundary และ owner
Contract
กำหนด schema, tool, permission, budget และ stop
Coordinate
เลือก parallel/sequential และรวมผลด้วย rubric
Recover
เก็บ checkpoint, trace, retry, resume และ escalation
สิ่งที่ต้องออกแบบ
checkpoint
ระบุ checkpoint พร้อมหลักฐานและเงื่อนไขล้มเหลว
ความซับซ้อนต้องซื้อความเร็วหรือคุณภาพที่วัดได้
resume
ระบุ resume พร้อมหลักฐานและเงื่อนไขล้มเหลว
checkpoint เก็บ completed units, artifacts, decisions, source versions และ next action; resume ต้อง validate state และไม่ทำผลข้างเคียงซ้ำ
idempotency
ระบุ idempotency พร้อมหลักฐานและเงื่อนไขล้มเหลว
checkpoint เก็บ completed units, artifacts, decisions, source versions และ next action; resume ต้อง validate state และไม่ทำผลข้างเคียงซ้ำ
version
ระบุ version พร้อมหลักฐานและเงื่อนไขล้มเหลว
checkpoint เก็บ completed units, artifacts, decisions, source versions และ next action; resume ต้อง validate state และไม่ทำผลข้างเคียงซ้ำ
ความซับซ้อนที่ไม่สร้างคุณค่า
แบ่งบทบาทจากชื่อตำแหน่ง
งานซ้ำและขอบเขตไม่ชัด
แบ่งจาก input/output/dependency
ให้ Agent แชร์การเขียน
race และข้อมูลทับกัน
single writer หรือ transaction
resume จากข้อความสุดท้าย
ทำ action ซ้ำ
durable checkpoint และ idempotency
ออกแบบงานหลายขั้น
งานหยุดหลังส่งอีเมลบางรายการแต่ก่อนบันทึกสถานะครบ
กำหนด decomposition, contracts, schedule, checkpoint, budget, approval และ trace
ดูคำตอบตัวอย่าง
เริ่มจาก DAG ของหน่วยงานอิสระ ใช้ parallel เฉพาะ read-only กำหนด JSON contract และ single writer บันทึก checkpoint ต่อหน่วยพร้อม source version ใช้ idempotency key กำหนด deadline/cost budget รวมผลด้วย rubric และให้ human approval ก่อนผลกระทบภายนอก
- DAG/dependency
- contract
- parallel/sequential
- checkpoint/resume
- budget/stop
- trace/approval
บันทึกช่วยจำ
- หลาย Agent เพิ่ม latency, cost และ failure surfaces ต้องพิสูจน์ประโยชน์ด้วย evaluation
- จำกัด credential และข้อมูลต่อบทบาท
สรุปบทเรียนนี้
- แยกงานตาม dependency
- ทำ contract และ single writer
- checkpoint trace และ recover