← Engineering Workflow

ENGINEERING WORKFLOW · 71

Deployment Incident:從 Failure 到 Success,怎麼不靠猜修 Production

真正的 incident debugging 不會一開始就知道答案。每一輪只根據當下 evidence 更新 hypothesis:最後成功在哪、第一個失敗在哪、下一個最低成本驗證是什麼。

Learning outcomes

1. Incident timeline

T0 commit pushed
T1 CI PASS
T2 deploy command PASS
T3 production smoke FAIL
T4 hypothesis tested
T5 fix deployed
T6 smoke PASS

時間線把「真的已驗證」和「我們猜應該正常」分開。

2. Deploy PASS 不等於 Production PASS

build artifact
  ↓
upload accepted
  ↓
public routing
  ↓
HTTP request
  ↓
content marker

Deployment tool 成功只證明前幾層。User 能否拿到正確內容,要靠 public smoke。

3. curl exit code 只是入口

curl -i -L   https://example.workers.dev/catalog

exit code 22 類似資訊只告訴你 HTTP status 被視為失敗;root cause 還要看 status、redirect chain、final URL、response body。

4. Redirect / canonical route

Static asset platform 可能把 /page.html 導向 /page。Smoke script 若沒 follow redirect,可能把正常服務誤判為 failure。

5. Fix 應改善未來 evidence

好的 incident fix 不只修這次 route,還會更新 smoke checks、logging、artifact checks,讓同類問題之後更早被發現。

Project checkpoint:Course Workspace v10

Commit SHA:
Workflow run:
Deploy version:
First failure:
Evidence:
Hypothesis:
Fix:
Regression check:

建立一份可交給別人重播的 incident report,而不是只留下「後來好了」。

Debug evidence:錯誤碼不是 root cause

同一個 HTTP failure 可能來自 routing、auth、platform config、artifact 缺檔等多個原因。先打開原始 evidence,再縮 hypothesis。

Knowledge check

  1. Deploy command PASS 能證明到哪?
  2. curl exit code 能否直接告訴 root cause?
  3. Incident timeline 的價值?
  4. 替 production 404 寫三個 competing hypotheses 與驗證方法。