Improve productivity with just these 5 engineering metrics

Five engineering metrics with precise definitions, formulas, realistic targets, and the pitfalls that make most teams measure them wrong.

BuildPulse Team

October 2, 2025

Improve productivity with just these 5 metrics | BuildPulse Blog

Every engineering metrics tool will happily show you forty charts. I have never seen a team improve because of chart thirty-one. The teams that get faster look at a handful of numbers, understand exactly how each one is computed, and know what to do when one of them moves.

This post is that handful. Five metrics, each with a precise definition, the formula (which events start and stop the clock), a realistic target, and the pitfall that makes most teams measure it wrong. Four of them are the stages of a pull request's life; the fifth is their sum. If you are still deciding whether to measure at all, start with why start measuring engineering metrics today. This is the post for after you have decided.

One rule before the list: all five are team-level metrics measured as medians (or percentiles) over a rolling window, not per-person and not as averages. Averages get dragged around by the one PR that sat open for a month. Medians tell you what a typical change experiences.

1. Coding time

Definition: how long a change is in active development before it is ready for someone else to look at.

Formula: time from the first commit on the branch to the moment the pull request is opened for review (leaving draft state, if you use drafts).

coding_time = pr.ready_for_review_at - branch.first_commit_at

What it tells you: how big your units of work are. Long coding time almost always means large changes, and large changes are slow at every subsequent stage too: harder to review, riskier to deploy, more likely to conflict.

A realistic target: a median under one working day. Teams that ship small, frequent changes routinely see medians of a few hours. If your median is measured in days, the problem is usually batch size rather than typing speed.

Pitfall: the first commit is a noisy start signal. Engineers rebase, squash, and open branches from stale work. Prefer the first commit that actually lands in the PR, and if your team squashes on merge, capture the original commit history before it disappears. Also, do not read coding time as "effort." A two-hour coding time on a change that took three days of thinking is not a two-hour change. This metric is about flow, not labor.

2. Pickup time

Definition: how long a pull request waits for a human to engage with it.

Formula: time from the PR being ready for review to the first review activity by someone other than the author (a comment, an approval, or a requested change).

pickup_time = first_reviewer_activity_at - pr.ready_for_review_at

What it tells you: whether review is a queue or a conversation. This is, in my experience, the single most common bottleneck in mid-sized engineering organizations, and the cheapest to fix, because it is almost pure waiting. Nobody is working on the PR during pickup time. It is just sitting there.

A realistic target: a median under four working hours, with a hard look at anything over a day. Some teams that treat review as an interrupt-level priority get to under an hour. Note the "working hours": a PR opened at 6pm Friday should not count the weekend against the reviewer.

Pitfall: bots. If a linter or coverage bot comments on every PR within thirty seconds, and you count that as "first review activity," your pickup time is thirty seconds forever and tells you nothing. Filter to human reviewers. The second pitfall is the author self-commenting; exclude the author too.

3. Review time

Definition: how long the review conversation takes once it has started.

Formula: time from the first reviewer activity to the final approval that unblocks the merge.

review_time = final_approval_at - first_reviewer_activity_at

What it tells you: how much back-and-forth a typical change generates, and whether your reviewers are reviewing or re-designing. Long review time with few comments means reviewers are slow to return. Long review time with many rounds means changes are arriving in a state that needs a lot of work, or the review culture rewards nitpicks.

A realistic target: a median under one working day. Short review time is not the goal by itself; a rubber-stamp culture gets a two-minute review time and a terrible change failure rate. Read this one alongside metric five.

Pitfall: "final approval" is fuzzy when you require multiple approvers, or when a late push dismisses an earlier approval. Decide on one rule (I use the approval after which no further changes were requested) and apply it consistently. The trend matters more than the definition, as long as the definition does not change under you. The deeper problem is that review time bundles two very different things, reviewer latency and change quality, and you will sometimes need to split them by looking at rounds per PR.

4. Deploy time

Definition: how long an approved change waits to reach production.

Formula: time from final approval to the deployment that includes the change. If merge and deploy are separate steps in your process, it is worth capturing both: approval to merge, and merge to deploy.

deploy_time = deployed_at - final_approval_at

What it tells you: how much of your pipeline is process rather than engineering. Release trains, manual QA gates, change-advisory boards, and slow CI all live here. In compliance-conscious environments some of this wait is intentional and correct; the point is to know how much of it there is and whether it is growing.

A realistic target: for teams with continuous deployment, under a few hours. For teams on a release cadence, the target is whatever your cadence implies (a daily train means a median around half a day), and the thing to watch is variance rather than the median.

Pitfall: CI reruns. If your test suite is flaky, a change that was approved at 10am and merged at 3pm may have spent those five hours waiting for a rerun to come back green, not waiting for a person. That is real time lost, and it belongs in deploy time, but it points to a different fix. Track your test suite's flaky failure rate next to this metric or you will misattribute the delay. We measured how much rerun time a typical suite burns in how much CI time flaky test reruns actually burn. The related question of what "merged" should mean is covered in what is merge time?.

5. Cycle time

Definition: the total elapsed time for a change from the first commit to production.

Formula: the sum of the four stages above, or equivalently:

cycle_time = deployed_at - branch.first_commit_at
           = coding_time + pickup_time + review_time + deploy_time

What it tells you: the one number to report upward. It is the engineering side of "how long does it take us to get an idea into customers' hands," and it is the metric most closely related to DORA's lead time for changes. DORA's published research has consistently associated shorter lead times with better organizational outcomes, and their "elite" band for lead time has been under a day in recent State of DevOps reports.

A realistic target: under two working days at the median for teams practicing small changes and continuous delivery. Under one day is achievable and worth aiming at. If you are at a week, do not panic and do not set a goal of one day; decompose it (that is what the other four metrics are for) and attack the biggest stage.

Pitfall: reporting cycle time without its stages. The number by itself only tells you that something is slow. The stages tell you what. We wrote about this decomposition in more depth in cycle time: the DORA metric that actually changes behavior. The second pitfall is measuring only merged PRs. Abandoned PRs are a signal too, and a team with a healthy cycle time but a 30% abandonment rate is not healthy.

Computing them from data you already have

You do not need a vendor to get a first baseline. Every event in the formulas above is in your code host's API. Here is a rough pickup-time calculation from the GitHub CLI, filtering out bots and the author:

# Median pickup time (hours) across the last 100 merged PRs
repo=your-org/your-repo
gh pr list --repo "$repo" --state merged --limit 100 --json number,author,createdAt \
  | jq -c '.[]' | while read -r pr; do
    num=$(echo "$pr" | jq -r .number)
    author=$(echo "$pr" | jq -r .author.login)
    opened=$(echo "$pr" | jq -r .createdAt)
    first=$(gh api "repos/$repo/pulls/$num/reviews" --paginate \
      | jq -r --arg a "$author" \
          '[.[] | select(.user.login != $a and (.user.type != "Bot"))] | min_by(.submitted_at) | .submitted_at // empty')
    [ -n "$first" ] && echo "$(( ($(date -d "$first" +%s) - $(date -d "$opened" +%s)) / 3600 ))"
  done | sort -n | awk '{a[NR]=$1} END {print "median pickup (hours):", a[int((NR+1)/2)]}'

It is slow, it ignores draft state and weekends, and it uses PR open rather than ready-for-review. All of that is fine for a first number. The point is to see the shape of your data before you commit to a definition.

Where a script stops being enough is the weekly, per-team, stage-decomposed version, with human-only reviewers, working-hours adjustment, and flaky CI failures separated from real ones. That is the maintenance burden our engineering metrics product exists to remove: it reads the same PR and CI events and keeps the five numbers current per team and repository, so the conversation can be about the numbers rather than the spreadsheet.

How the five fit together

Read them as a chain, because that is what they are. A long cycle time is a symptom. The four stages are the diagnosis. Each stage has a characteristic fix:

  • Long coding time: smaller changes, feature flags, stacked PRs.
  • Long pickup time: review ownership, notifications, a norm that review comes before new work.
  • Long review time: better PR descriptions, smaller diffs, agreed standards so reviews are checks rather than debates.
  • Long deploy time: faster CI, fewer manual gates, and a test suite you can trust so nobody is waiting on reruns.

The one thing that undermines all five at once is a CI signal you cannot trust. If a meaningful fraction of runs fail on unchanged code, pickup time goes up (nobody reviews a red PR), review time goes up (rounds get wasted on "is this failure real?"), and deploy time goes up (reruns). Fix the signal first or every number you collect afterward will be partly measuring your flaky tests.

FAQ

Should I track these per engineer?

No, or at least not at first, and never for evaluation. Per-person cycle time is dominated by the type of work someone is assigned, not by how good they are. Team-level medians over a rolling four-week window tell you about the process, which is the thing you can actually change.

What about DORA's four metrics?

Cycle time is close to DORA's lead time for changes. Deploy frequency and change failure rate are outcome metrics worth adding once the five above are stable; they answer "how often and how safely" rather than "how long." Time to restore matters if you run production services. The five in this post are the stage-level diagnostics that tell you why a DORA number moved.

Which one should I fix first?

Whichever stage is the largest share of your cycle time. For most teams I have seen that is pickup time, and it is also the cheapest to improve because it is pure waiting. Measure all five, pick the biggest, change one thing, and check again in a month.

Stop guessing which tests you can trust

BuildPulse finds your flaky tests, ranks them by the engineering time they cost, and lets you quarantine the worst in one click. See results on your first build.

Free to start · No credit card required · Setup is a single CI step