A human engineer changes three files and writes a commit message that still means something in two weeks. A coding agent changes thirty, refactors a helper “while it is here,” upgrades a default, and leaves a paragraph of confidence.
CI is green. You merge.
Three days later POST /evaluate-credit fails on one partner payload and nobody can point at the line.
The table below shows how differently the two kinds of commit arrive in your history.
| Human engineer | Coding agent | |
| Files changed | 3 | 30, plus a helper refactored “while it is here” and an upgraded default |
| Commit message | Still means something in two weeks | A paragraph of confidence |
Regressions From Coding Agents Are a Search Problem
That is not a mystery. That is a search problem. Git solved search problems in 2007.
git bisect is binary search over history.
You give it one commit that is good and one that is bad. It checks out the midpoint. You tell it the truth, or a script does. It throws away half the range. After log2(n) steps you are not guessing which agent turn went feral. You are looking at the first bad commit.
200 commits is eight questions. Not a war room.
The agentic era did not invent regressions. It made them cheap, frequent, and hard to smell. It also gave us something 2007 did not have: an agent that can write the predicate, run git bisect run, and stop talking the moment Git prints is the first bad commit.
That is the enterprise move. Use the model to drive the loop. Do not use the model to invent a villain.
1. The System Under Test: A Three-Agent Credit Screen
The service is a three-agent credit screen behind one boring HTTP contract.

Each agent has one job, and each one has a failure mode coding agents love to introduce.
| Agent | Job | Failure mode agents love to introduce |
| Financial | Parse raw_text for monthly income, expenses, cashflow | Drop a locale scale (lakh, crore) during a “cleanup” |
| Risk | DTI = existing_debts / monthly_income, band at 0.40 / 0.20 | Rewrite the formula, keep the tests that never hit the band |
| Managerial | Call both, apply policy | Patch _decide because the model guessed the symptom |
The API contract
The service exposes two routes.
| Method | Route | Purpose |
| GET | /health | Liveness check |
| POST | /evaluate-credit | Runs the full evaluation and returns an EvaluationResult |
The request carries free text, plus optional debts and a name.
| CreditRequest field | Type |
| raw_text | str (required) |
| existing_debts | Optional[float], defaults to None |
| applicant_name | Optional[str], defaults to None |
The response carries the financial metrics, the risk assessment, and the decision.
| EvaluationResult field | Type | Contains |
| applicant_name | Optional[str] | Name from the request |
| financial | FinancialMetrics | income, expenses, cashflow, status |
| risk | RiskAssessment | debts, dti, risk_level, rationale |
| decision | Decision | approve, review, or reject |
| decision_reason | str | Why the decision was made |
| confidence | str | Confidence of the decision |
The endpoint awaits orchestrator.run_evaluation(request). Any exception becomes an HTTP 500 with the detail Evaluation failed: {exc}.
The manager composes, it does not think
run_evaluation does no reasoning of its own. It runs four steps in order:
- Calls financial_agent.extract_metrics(request.raw_text).
- Parses amounts from the same text and takes the largest debt found.
- Uses request.existing_debts if it was supplied, otherwise the debt parsed from the text, and passes it with monthly_income to risk_agent.analyze_risk.
- Calls self._decide(financial, risk) for the decision, reason, and confidence, and returns the EvaluationResult.
The risk agent: three numbers and a sentence
Risk is three numbers and a sentence. That is the kind of code agents “improve.” DTI is round(debts / monthly_income, 4), then banded:
| DTI | Risk level |
| >= 0.40 (HIGH_DTI) | RiskLevel.HIGH |
| >= 0.20 (MEDIUM_DTI) | RiskLevel.MEDIUM |
| Below 0.20 | RiskLevel.LOW |
The load-bearing surface: turning “1.2 lakh” into a float
The load-bearing surface is not policy. It is how a sentence becomes a float in India.

An applicant does not type 120000. They type 1.2 lakh. If that key dies, monthly_income is null and the manager does the only honest thing left: send the case to review.

It looks like the model failed. The history failed.
2. Why Agents Make Bisect Mandatory, and Why They Also Make It Cheap
Two facts, same era. The table below sets them side by side.
| Fact | What it looks like in practice |
| Agents raise commit velocity | A single session can land twenty hashes. Messages say chore: normalise amount handling. Tests assert the happy path the agent itself invented (80,000, never 1.2 lakh). Review skims intent. Surface area is what actually moved. |
| Agents can run a binary search without an opinion | The model is a bad archaeologist and a good subprocess. Constrained correctly, it writes scripts/repro.sh, calls git bisect run, and reads only the culprit diff. Unconstrained, it runs git log, hallucinates a cause, and patches _decide. |

Enterprise rule: the agent is allowed to operate Git. It is not allowed to skip Git.
The prompt that actually works in a coding-agent harness
Constraint: use ONLY git bisect to locate the culprit. Do not use git blame, git log archaeology, or a guess from the diff.
- Identify a last-known-good tag.
- Write scripts/repro.sh: exit 0 = property holds (good), exit 1 = property failed (bad), exit 125 = commit cannot be tested (skip).
- Run git bisect start, git bisect bad HEAD, git bisect good <tag>, then git bisect run ./scripts/repro.sh.
- Stop when Git prints <hash> is the first bad commit.
- Run git show <hash>. Explain that diff only. Propose revert or patch.
- Run git bisect reset.
That is the agentic contribution. Not “AI found the bug.” The harness found the commit. The model was a pair of hands.
3. The Full git bisect Command Surface
Do not memorize half of this and improvise the rest on an incident. This is the whole tool.
3.1 Start a session and mark the endpoints
Start a session step by step, or as a one-liner with the bad commit first and one or more good commits after it.
| Command | What it does |
| git bisect help | Shows help |
| git bisect start | Starts a session (current HEAD is usually the bad end) |
| git bisect start <bad> <good> [<good>…] [–] [<pathspec>…] | One-liner: bad first, then one or more good commits |
| git bisect start HEAD v1.4.2 | Example: HEAD is bad, v1.4.2 is good |
| git bisect start HEAD v1.4.2 v1.4.0 | Multiple known-good tips |
| git bisect start HEAD v1.4.2 — evaluate_credit.py | Only this path |
Then mark the endpoints.
| Command | What it marks |
| git bisect bad | Current checkout is bad |
| git bisect bad HEAD / git bisect bad abc1234 | A named commit is bad |
| git bisect good | Current checkout is good |
| git bisect good v1.4.2 | A named commit is good |
| git bisect good abc0000 def1111 | Several goods |
Choose the right vocabulary
“Good” and “bad” are not always the right metaphor. Old aliases carry the same meaning, and custom terms cover properties that are not bugs.
| When you are hunting | Newer end (bad) | Older end (good) | How to set it |
| A regression | bad | good | Default |
| Performance or feature flags | new | old | Built-in synonyms for bad and good |
| A property that is not a bug | slow | fast | git bisect start –term-old fast –term-new slow |
| The commit that fixed something | fixed | broken | git bisect start –term-new fixed –term-old broken |
With custom terms, git bisect terms prints the terms this session is using. You then mark with your own words: git bisect slow HEAD and git bisect fast v1.4.2, or git bisect broken v1.4.2 and git bisect fixed HEAD.
3.2 Walk by hand
After start, Git checks out a midpoint and prints where you are:
Bisecting: 23 revisions left to test after this (roughly 5 steps) [a3f1d22] refactor: tighten types on CreditRequest
You reproduce, then report one verdict per midpoint.
| Command | Use it when |
| git bisect good | The property holds here |
| git bisect bad | The property failed here |
| git bisect skip | This commit cannot be tested |
| git bisect skip HEAD~3..HEAD | A whole range cannot be tested (broken build island) |
skip is not “I am tired.” It is “this revision does not compile / cannot boot / is unrelated merge noise.” Abuse it and the first-bad-commit answer becomes a range instead of a point.
3.3 Walk unattended with git bisect run
This is the command that makes agents useful: git bisect run <cmd> [<arg>…].
The exit-code contract decides everything. Memorize it. The agent’s script must obey it.
| Exit code | Meaning | Git’s action |
| 0 | good / old / property holds | cut the right half |
| 1–124, 126–127 | bad / new / property failed | cut the left half |
| 125 | cannot test this revision | same as git bisect skip |
| -1 / 255 / anything else | script crashed | abort the session |

Any command that honors the contract works.
| Command | Runs |
| git bisect run ./scripts/repro.sh | A repro script |
| git bisect run make test | A make target |
| git bisect run pytest tests/test_locale_amounts.py -q | A single test file |
| git bisect run –reset-when-found ./scripts/repro.sh | Returns you to the starting branch as soon as the culprit is identified (available on current Git) |
3.4 Limit the search and keep merge noise out
Three options shrink the search before it starts.
| Option | Example | What it does | When to use it |
| — <pathspec> | git bisect start HEAD v1.4.2 — evaluate_credit.py scripts/ | Only commits that touched the parser / API | Stops Git from dragging you through unrelated docs and lockfile churn |
| –first-parent | git bisect start –first-parent HEAD v1.4.2 | Ignores merged topic-branch interiors | The enterprise default on repos that squash-merge poorly and merge often |
| –no-checkout | git bisect start –no-checkout HEAD v1.4.2 | Does not check out a working tree; operates by SHA | Bare repos, very large trees, scripted agents |
3.5 Inspect, record, replay, and leave
The housekeeping commands make a session auditable and transferable.
| Job | Command |
| See every good/bad/skip decision this session | git bisect log |
| Save the session | git bisect log > /tmp/bisect.log |
| Rebuild the same session later (handoff between on-call and the next shift) | git bisect replay /tmp/bisect.log |
| See remaining candidates (gitk) | git bisect visualize or git bisect view |
| Same, if visualize is aliased to git log | git bisect view –oneline |
| Return to the branch you started on | git bisect reset |
| Return to a named ref instead | git bisect reset main |
After Git names the commit, run git show <first-bad>, then git revert <first-bad> when the commit is atomic, then git bisect reset.
If the agent squashed six ideas into one hash, revert is a blunt instrument. That is a process failure that happened before the incident, not a bisect failure.
4. Worked Session: Bisecting /evaluate-credit and “1.2 lakh”
4.1 Pin two payloads
A bisect needs two payloads: one that proves the predicate is honest, and one that reproduces the break. Both are sent as a POST to http://127.0.0.1:8000/evaluate-credit with Content-Type: application/json.
| Payload | applicant_name | raw_text | existing_debts | Role |
| Control | Asha | “I earn 80,000 rupees a month. Rent and bills are about 25,000. Outstanding loans total 180,000.” | Not sent | Must keep passing on every midpoint, or your predicate is lying |
| Repro | Meera | “Monthly income 1.2 lakh. Bills 20k.” | 50000 | The partner body that started returning review-with-null-income |
When the code is good, the repro payload returns this property:
| Field | Expected value |
| financial.monthly_income | 120000.0 |
| financial.monthly_expenses | 20000.0 |
| financial.financial_status | surplus |
| risk.existing_debts | 50000.0 |
| risk.dti | 0.4167 |
| risk.risk_level | high |
| decision | reject |
When the scale table has lost lakh, monthly_income is null and the manager returns review. That flip is the predicate. Not “the page feels wrong.”
4.2 Write a predicate the agent is allowed to run
scripts/repro.sh is executable, idempotent, and uses no network besides localhost. It skips commits that cannot import the stack, boots the service, posts the repro payload, and checks the income and DTI.

Make it executable with chmod +x scripts/repro.sh.
4.3 Run it
Four commands start the search and hand it to the script:
- git bisect start
- git bisect bad HEAD
- git bisect good v1.4.2
- git bisect run ./scripts/repro.sh
The table below shows what the walk looked like, one verdict per step.
| Step | Commit | repro.sh |
| 1 | a3f1d22 types on CreditRequest | 0: 1.2 lakh → 120000 |
| 2 | b7c91aa extract-metrics delay | 0 |
| 3 | c12e08f rewrite _decide copy | 0 |
| 4 | d9ab440 chore: normalise amount handling | 1: income null |
| 5 | parent of d9ab440 | 0 |

Git’s only sentence that matters: d9ab440c1e8f is the first bad commit
| Detail | Value |
| Commit | d9ab440c1e8f |
| Author | coding-agent <agent@local> |
| Date | Thu Sep 18 22:11:03 2026 +0530 |
| Message | chore: normalise amount handling |
| Change | evaluate_credit.py: 1 file changed, 8 insertions(+), 10 deletions(-) |
4.4 Read the diff, revert, reset
Run git show d9ab440c1e8f.
_AMOUNT_PATTERN still captured the word lakh. _to_float no longer multiplied it. The unit test used 80000. Suite green. Partner dead.
Run git revert d9ab440c1e8f, then git bisect reset.
Two minutes to see it. One revert. Postmortem: we reviewed intent and not surface area.
If you must hand the session to the next on-call instead of finishing it, save it with git bisect log > /tmp/evaluate-credit.bisect.log, then, later on another clone, rebuild it with git bisect replay /tmp/evaluate-credit.bisect.log.
5. How a Senior Team Actually Staffs This
Six operating rules keep bisect fast when agents write the commits. The table below summarizes them; each one is expanded underneath.
| Rule | In one line |
| Predicate before theory | Name a testable property before you name a cause |
| Atomic commits, even when an agent wrote them | Unrelated changes never share a hash |
| CI owns a locale fixture | 80,000 is necessary, not sufficient |
| 125 is a first-class answer | Untestable midpoints skip; they are never marked bad |
| Agents get a closed loop, not an open prompt | Write / run / show / reset |
| Pathspec when you already know the subsystem | Start narrow, widen only if the file never moved |
Predicate before theory
“p95 of /evaluate-credit exceeds 200ms” is a predicate. “DTI is null for input X” is a predicate. “The agents feel off” is not.
- Atomic commits, even when an agent wrote them
extract_metrics and “rewrite DTI bands” never share a hash. Squash-merging a 40-file agent session into wip agent changes makes bisect a brick. If you must squash for product branches, keep the agent’s working branch unsquashed long enough to bisect.
- CI owns a locale fixture
The control payload (80,000) is necessary. It is not sufficient. Ship 1.2 lakh, 20k, and 1 crore in the same job that gates merge. The agent will not invent those cases. You will.
- 125 is a first-class answer
Midpoints that cannot import FastAPI, cannot bind the port, or predate the endpoint should skip. Do not mark them bad. You will pin the wrong commit.
- Agents get a closed loop, not an open prompt
Write / run / show / reset. No git log tourism. No speculative patch to _decide before the hash exists.
- Pathspec when you already know the subsystem
Credit regressions start with git bisect start HEAD v1.4.2 — evaluate_credit.py. Widen only if that search says the file never moved.
6. git bisect Command Cheat Sheet
Every command from this guide, in one place.

And the exit-code contract, one last time:
| Exit code | Result |
| 0 | good |
| 1 | bad |
| 125 | skip |
| other | abort |
The Era Changed the Speed of the Mistake, Not the Maths
Use agents to write the three specialists. Use agents to write repro.sh and run git bisect run. Use bisect to keep a spine in the history so that when the helpful machine gets clever with a regex, you can find the exact second it did.
The era changed the speed of the mistake. It did not change the maths.
Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here
