All posts
AI Agents

We Reviewed the PR. Bisect Reviewed Time.

Lyzr Team
Lyzr Team
Oct 1, 2026
12 min read
We Reviewed the PR. Bisect Reviewed Time.

A human engineer changes three files and writes a commit message that still means something in two weeks. A coding agent changes thirty, refactors a helper “while it is here,” upgrades a default, and leaves a paragraph of confidence.

CI is green. You merge. 

Three days later POST /evaluate-credit fails on one partner payload and nobody can point at the line.

The table below shows how differently the two kinds of commit arrive in your history.

Human engineerCoding agent
Files changed330, plus a helper refactored “while it is here” and an upgraded default
Commit messageStill means something in two weeksA paragraph of confidence

Regressions From Coding Agents Are a Search Problem

That is not a mystery. That is a search problem. Git solved search problems in 2007.

git bisect is binary search over history. 

You give it one commit that is good and one that is bad. It checks out the midpoint. You tell it the truth, or a script does. It throws away half the range. After log2(n) steps you are not guessing which agent turn went feral. You are looking at the first bad commit.

200 commits is eight questions. Not a war room.

The agentic era did not invent regressions. It made them cheap, frequent, and hard to smell. It also gave us something 2007 did not have: an agent that can write the predicate, run git bisect run, and stop talking the moment Git prints is the first bad commit.

That is the enterprise move. Use the model to drive the loop. Do not use the model to invent a villain.

1. The System Under Test: A Three-Agent Credit Screen

The service is a three-agent credit screen behind one boring HTTP contract.

image 2
We Reviewed the PR. Bisect Reviewed Time. 9

Each agent has one job, and each one has a failure mode coding agents love to introduce.

AgentJobFailure mode agents love to introduce
FinancialParse raw_text for monthly income, expenses, cashflowDrop a locale scale (lakh, crore) during a “cleanup”
RiskDTI = existing_debts / monthly_income, band at 0.40 / 0.20Rewrite the formula, keep the tests that never hit the band
ManagerialCall both, apply policyPatch _decide because the model guessed the symptom

The API contract

The service exposes two routes.

MethodRoutePurpose
GET/healthLiveness check
POST/evaluate-creditRuns the full evaluation and returns an EvaluationResult

The request carries free text, plus optional debts and a name.

CreditRequest fieldType
raw_textstr (required)
existing_debtsOptional[float], defaults to None
applicant_nameOptional[str], defaults to None

The response carries the financial metrics, the risk assessment, and the decision.

EvaluationResult fieldTypeContains
applicant_nameOptional[str]Name from the request
financialFinancialMetricsincome, expenses, cashflow, status
riskRiskAssessmentdebts, dti, risk_level, rationale
decisionDecisionapprove, review, or reject
decision_reasonstrWhy the decision was made
confidencestrConfidence of the decision

The endpoint awaits orchestrator.run_evaluation(request). Any exception becomes an HTTP 500 with the detail Evaluation failed: {exc}.

The manager composes, it does not think

run_evaluation does no reasoning of its own. It runs four steps in order:

  1. Calls financial_agent.extract_metrics(request.raw_text).
  2. Parses amounts from the same text and takes the largest debt found.
  3. Uses request.existing_debts if it was supplied, otherwise the debt parsed from the text, and passes it with monthly_income to risk_agent.analyze_risk.
  4. Calls self._decide(financial, risk) for the decision, reason, and confidence, and returns the EvaluationResult.

The risk agent: three numbers and a sentence

Risk is three numbers and a sentence. That is the kind of code agents “improve.” DTI is round(debts / monthly_income, 4), then banded:

DTIRisk level
>= 0.40 (HIGH_DTI)RiskLevel.HIGH
>= 0.20 (MEDIUM_DTI)RiskLevel.MEDIUM
Below 0.20RiskLevel.LOW

The load-bearing surface: turning “1.2 lakh” into a float

The load-bearing surface is not policy. It is how a sentence becomes a float in India.

image 5
We Reviewed the PR. Bisect Reviewed Time. 10

An applicant does not type 120000. They type 1.2 lakh. If that key dies, monthly_income is null and the manager does the only honest thing left: send the case to review.

image 4
We Reviewed the PR. Bisect Reviewed Time. 11

It looks like the model failed. The history failed.

2. Why Agents Make Bisect Mandatory, and Why They Also Make It Cheap

Two facts, same era. The table below sets them side by side.

FactWhat it looks like in practice
Agents raise commit velocityA single session can land twenty hashes. Messages say chore: normalise amount handling. Tests assert the happy path the agent itself invented (80,000, never 1.2 lakh). Review skims intent. Surface area is what actually moved.
Agents can run a binary search without an opinionThe model is a bad archaeologist and a good subprocess. Constrained correctly, it writes scripts/repro.sh, calls git bisect run, and reads only the culprit diff. Unconstrained, it runs git log, hallucinates a cause, and patches _decide.
image
We Reviewed the PR. Bisect Reviewed Time. 12

Enterprise rule: the agent is allowed to operate Git. It is not allowed to skip Git.

The prompt that actually works in a coding-agent harness

Constraint: use ONLY git bisect to locate the culprit. Do not use git blame, git log archaeology, or a guess from the diff.

  1. Identify a last-known-good tag.
  2. Write scripts/repro.sh: exit 0 = property holds (good), exit 1 = property failed (bad), exit 125 = commit cannot be tested (skip).
  3. Run git bisect start, git bisect bad HEAD, git bisect good <tag>, then git bisect run ./scripts/repro.sh.
  4. Stop when Git prints <hash> is the first bad commit.
  5. Run git show <hash>. Explain that diff only. Propose revert or patch.
  6. Run git bisect reset.

That is the agentic contribution. Not “AI found the bug.” The harness found the commit. The model was a pair of hands.

3. The Full git bisect Command Surface

Do not memorize half of this and improvise the rest on an incident. This is the whole tool.

3.1 Start a session and mark the endpoints

Start a session step by step, or as a one-liner with the bad commit first and one or more good commits after it.

CommandWhat it does
git bisect helpShows help
git bisect startStarts a session (current HEAD is usually the bad end)
git bisect start <bad> <good> [<good>…] [–] [<pathspec>…]One-liner: bad first, then one or more good commits
git bisect start HEAD v1.4.2Example: HEAD is bad, v1.4.2 is good
git bisect start HEAD v1.4.2 v1.4.0Multiple known-good tips
git bisect start HEAD v1.4.2 — evaluate_credit.pyOnly this path

Then mark the endpoints.

CommandWhat it marks
git bisect badCurrent checkout is bad
git bisect bad HEAD / git bisect bad abc1234A named commit is bad
git bisect goodCurrent checkout is good
git bisect good v1.4.2A named commit is good
git bisect good abc0000 def1111Several goods

Choose the right vocabulary

“Good” and “bad” are not always the right metaphor. Old aliases carry the same meaning, and custom terms cover properties that are not bugs.

When you are huntingNewer end (bad)Older end (good)How to set it
A regressionbadgoodDefault
Performance or feature flagsnewoldBuilt-in synonyms for bad and good
A property that is not a bugslowfastgit bisect start –term-old fast –term-new slow
The commit that fixed somethingfixedbrokengit bisect start –term-new fixed –term-old broken

With custom terms, git bisect terms prints the terms this session is using. You then mark with your own words: git bisect slow HEAD and git bisect fast v1.4.2, or git bisect broken v1.4.2 and git bisect fixed HEAD.

3.2 Walk by hand

After start, Git checks out a midpoint and prints where you are:

Bisecting: 23 revisions left to test after this (roughly 5 steps) [a3f1d22] refactor: tighten types on CreditRequest

You reproduce, then report one verdict per midpoint.

CommandUse it when
git bisect goodThe property holds here
git bisect badThe property failed here
git bisect skipThis commit cannot be tested
git bisect skip HEAD~3..HEADA whole range cannot be tested (broken build island)

skip is not “I am tired.” It is “this revision does not compile / cannot boot / is unrelated merge noise.” Abuse it and the first-bad-commit answer becomes a range instead of a point.

3.3 Walk unattended with git bisect run

This is the command that makes agents useful: git bisect run <cmd> [<arg>…].

The exit-code contract decides everything. Memorize it. The agent’s script must obey it.

Exit codeMeaningGit’s action
0good / old / property holdscut the right half
1–124, 126–127bad / new / property failedcut the left half
125cannot test this revisionsame as git bisect skip
-1 / 255 / anything elsescript crashedabort the session
image 6
We Reviewed the PR. Bisect Reviewed Time. 13

Any command that honors the contract works.

CommandRuns
git bisect run ./scripts/repro.shA repro script
git bisect run make testA make target
git bisect run pytest tests/test_locale_amounts.py -qA single test file
git bisect run –reset-when-found ./scripts/repro.shReturns you to the starting branch as soon as the culprit is identified (available on current Git)

3.4 Limit the search and keep merge noise out

Three options shrink the search before it starts.

OptionExampleWhat it doesWhen to use it
— <pathspec>git bisect start HEAD v1.4.2 — evaluate_credit.py scripts/Only commits that touched the parser / APIStops Git from dragging you through unrelated docs and lockfile churn
–first-parentgit bisect start –first-parent HEAD v1.4.2Ignores merged topic-branch interiorsThe enterprise default on repos that squash-merge poorly and merge often
–no-checkoutgit bisect start –no-checkout HEAD v1.4.2Does not check out a working tree; operates by SHABare repos, very large trees, scripted agents

3.5 Inspect, record, replay, and leave

The housekeeping commands make a session auditable and transferable.

JobCommand
See every good/bad/skip decision this sessiongit bisect log
Save the sessiongit bisect log > /tmp/bisect.log
Rebuild the same session later (handoff between on-call and the next shift)git bisect replay /tmp/bisect.log
See remaining candidates (gitk)git bisect visualize or git bisect view
Same, if visualize is aliased to git loggit bisect view –oneline
Return to the branch you started ongit bisect reset
Return to a named ref insteadgit bisect reset main

After Git names the commit, run git show <first-bad>, then git revert <first-bad> when the commit is atomic, then git bisect reset.

If the agent squashed six ideas into one hash, revert is a blunt instrument. That is a process failure that happened before the incident, not a bisect failure.

4. Worked Session: Bisecting /evaluate-credit and “1.2 lakh”

4.1 Pin two payloads

A bisect needs two payloads: one that proves the predicate is honest, and one that reproduces the break. Both are sent as a POST to http://127.0.0.1:8000/evaluate-credit with Content-Type: application/json.

Payloadapplicant_nameraw_textexisting_debtsRole
ControlAsha“I earn 80,000 rupees a month. Rent and bills are about 25,000. Outstanding loans total 180,000.”Not sentMust keep passing on every midpoint, or your predicate is lying
ReproMeera“Monthly income 1.2 lakh. Bills 20k.”50000The partner body that started returning review-with-null-income

When the code is good, the repro payload returns this property:

FieldExpected value
financial.monthly_income120000.0
financial.monthly_expenses20000.0
financial.financial_statussurplus
risk.existing_debts50000.0
risk.dti0.4167
risk.risk_levelhigh
decisionreject

When the scale table has lost lakh, monthly_income is null and the manager returns review. That flip is the predicate. Not “the page feels wrong.”

4.2 Write a predicate the agent is allowed to run

scripts/repro.sh is executable, idempotent, and uses no network besides localhost. It skips commits that cannot import the stack, boots the service, posts the repro payload, and checks the income and DTI.

image 7
We Reviewed the PR. Bisect Reviewed Time. 14

Make it executable with chmod +x scripts/repro.sh.

4.3 Run it

Four commands start the search and hand it to the script:

  1. git bisect start
  2. git bisect bad HEAD
  3. git bisect good v1.4.2
  4. git bisect run ./scripts/repro.sh

The table below shows what the walk looked like, one verdict per step.

StepCommitrepro.sh
1a3f1d22 types on CreditRequest0: 1.2 lakh → 120000
2b7c91aa extract-metrics delay0
3c12e08f rewrite _decide copy0
4d9ab440 chore: normalise amount handling1: income null
5parent of d9ab4400
image 1
We Reviewed the PR. Bisect Reviewed Time. 15

Git’s only sentence that matters: d9ab440c1e8f is the first bad commit

DetailValue
Commitd9ab440c1e8f
Authorcoding-agent <agent@local>
DateThu Sep 18 22:11:03 2026 +0530
Messagechore: normalise amount handling
Changeevaluate_credit.py: 1 file changed, 8 insertions(+), 10 deletions(-)

4.4 Read the diff, revert, reset

Run git show d9ab440c1e8f.

_AMOUNT_PATTERN still captured the word lakh. _to_float no longer multiplied it. The unit test used 80000. Suite green. Partner dead.

Run git revert d9ab440c1e8f, then git bisect reset.

Two minutes to see it. One revert. Postmortem: we reviewed intent and not surface area.

If you must hand the session to the next on-call instead of finishing it, save it with git bisect log > /tmp/evaluate-credit.bisect.log, then, later on another clone, rebuild it with git bisect replay /tmp/evaluate-credit.bisect.log.

5. How a Senior Team Actually Staffs This

Six operating rules keep bisect fast when agents write the commits. The table below summarizes them; each one is expanded underneath.

RuleIn one line
Predicate before theoryName a testable property before you name a cause
Atomic commits, even when an agent wrote themUnrelated changes never share a hash
CI owns a locale fixture80,000 is necessary, not sufficient
125 is a first-class answerUntestable midpoints skip; they are never marked bad
Agents get a closed loop, not an open promptWrite / run / show / reset
Pathspec when you already know the subsystemStart narrow, widen only if the file never moved

Predicate before theory

“p95 of /evaluate-credit exceeds 200ms” is a predicate. “DTI is null for input X” is a predicate. “The agents feel off” is not.

  1. Atomic commits, even when an agent wrote them

extract_metrics and “rewrite DTI bands” never share a hash. Squash-merging a 40-file agent session into wip agent changes makes bisect a brick. If you must squash for product branches, keep the agent’s working branch unsquashed long enough to bisect.

  1. CI owns a locale fixture

The control payload (80,000) is necessary. It is not sufficient. Ship 1.2 lakh, 20k, and 1 crore in the same job that gates merge. The agent will not invent those cases. You will.

  1. 125 is a first-class answer

Midpoints that cannot import FastAPI, cannot bind the port, or predate the endpoint should skip. Do not mark them bad. You will pin the wrong commit.

  1. Agents get a closed loop, not an open prompt

Write / run / show / reset. No git log tourism. No speculative patch to _decide before the hash exists.

  1. Pathspec when you already know the subsystem

Credit regressions start with git bisect start HEAD v1.4.2 — evaluate_credit.py. Widen only if that search says the file never moved.

6. git bisect Command Cheat Sheet

Every command from this guide, in one place.

image 3
We Reviewed the PR. Bisect Reviewed Time. 16

And the exit-code contract, one last time:

Exit codeResult
0good
1bad
125skip
otherabort

The Era Changed the Speed of the Mistake, Not the Maths

Use agents to write the three specialists. Use agents to write repro.sh and run git bisect run. Use bisect to keep a spine in the history so that when the helpful machine gets clever with a regex, you can find the exact second it did.

The era changed the speed of the mistake. It did not change the maths.

Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here
Build with Lyzr

Try it in
Agent Studio

From framework-agnostic design to production-grade agents, deployed in under 24 hours.