Been dealing with this more often lately. Tests pass on my machine, I push, and CI blows up. Usually it’s one of these:
- Different Node/Python/whatever version
- Missing env vars that exist in my .env but not in CI secrets
- File system case sensitivity (macOS vs Linux)
- Some flaky test that depends on timing
My current debugging flow is pretty basic: check the logs, compare versions, run the exact same Docker image locally if I can. But it still eats 20-30 minutes each time before I figure out the actual problem.
Anyone have a more systematic approach? Like a quick checklist you run through before you even look at the logs?
Also curious — do you replicate your CI environment locally with something like act (for GitHub Actions) or just trust the remote runner?


If something fails look at the logs, actually READ them, what does it sail failed? Investigate that.
Failure + no logs?
Your mission is now to first find a way to enable logging, then read them.
No way to enable logging? If the infra and code is yours, now you need to write in logging, enable it, and then read the log
Not your code? No logs?
Figure out the difference between the last working version of your code / deployment script and the current one. If no difference check the change log for the infra you are using
That’s my brief list