How to Measure Developer Experience (Without Turning It Into a Scorecard)
What I actually do when a team asks me to measure developer productivity—and why the survey matters more than the dashboard.
"Our builds take 45 minutes"
Someone once asked me: "Our builds take 45 minutes. How do we make them faster?"
I asked, "Why do they take 45 minutes?"
They looked at me like I'd grown a second head.
That's the whole problem with developer productivity metrics, in one conversation. Walk into any tech company right now and you'll find DORA on a dashboard, SPACE in the quarterly review, cycle time in every sprint retro. Everyone's measuring. Almost no one's understanding. Teams treat these numbers like scorecards instead of diagnostics—mostly because "that's what Google does" or "that's what the McKinsey report said."
The metrics aren't the villain here. DORA is great. SPACE is thoughtful. The problem is using them as endpoints instead of starting points. Metrics tell you where to look, not what to do.
So this is how I actually measure developer experience when a team brings me in. It's less about tooling than you'd think.
What you're actually measuring
Developer experience is how it feels to get work done at your company. How long you sit there waiting for a build. How much you have to hold in your head to ship one change. Whether you ever get three hours to yourself.
The best breakdown I've seen is from a 2023 paper by Abi Noda, Margaret-Anne Storey, Nicole Forsgren, and Michaela Greiler. They took 25-plus factors and boiled them down to three:
Feedback loops. How fast do you find out whether the thing you did worked? That's your build, your tests, your CI—and it's people, too. How long until someone reviews your PR, or answers your question in chat. A slow loop costs more than the waiting. People batch up changes, go do something else, and forget where they were. That's where the hours go.
Cognitive load. How much do you have to know to get one task done? At 8x8 we had hundreds of APIs that looked like they came from a dozen different companies, and documentation spread out across 7+ places. Our engineers weren't slow. Every task just started with figuring out which of the seven docs was the right one.
Flow. How often do you get a real stretch of focus? Meetings, on-call noise, unplanned work, chat pings—they chop the day up. The hard problems get solved in flow. Great tooling and no focus time still ships slow.
Every number on this page maps to one of those three. If a metric doesn't, I don't collect it. The same paper suggests measuring each one two ways—ask people, and pull from systems:
| Dimension | Ask people | Pull from systems |
|---|---|---|
| Feedback loops | How happy are you with build/test speed? With code review speed? | CI duration, time to first review, PR pickup time |
| Cognitive load | How complex does the codebase feel? How easy is it to find an answer? | Time to get a technical question answered; new hire time to first merged change |
| Flow | Can you focus? Do you get enough uninterrupted time? | Meeting-free blocks per week; how often unplanned work shows up |
Plus a few overall numbers: how easy people say it is to ship, how productive they feel, whether they're engaged. That's your "is any of this working?" check.
The numbers, if you need to convince someone
If you've got to make the case to a CFO or a skeptical VP, you don't have to wing it. The research has gotten pretty good over the last five years.
69%
of developers lose 8+ hours a week to friction
About a day out of every five. Tech debt, bad docs, slow builds, no focus time. Fewer than half of them think their leaders know. (Atlassian, 2024)
50%
more productive with real time for deep work
Intuitive tools: 50% more innovative. Tight feedback loops: 50% less tech debt reported. (Forsgren, Noda et al., "DevEx in Action," 2024)
82% vs 7%
odds of a "good day" with few interruptions vs. many
One meeting a day: 99% chance of a good day. Three meetings: 14%. (GitHub's Good Day Project, 2021)
62%
say tech debt is their biggest frustration
Double the next item on the list. More than 60% spend 30+ minutes a day hunting for answers. (Stack Overflow Developer Survey, 2024)
The one that surprised me the first time I read it: Google surveyed 622 developers across three companies to see what predicts how productive people feel. The top three had nothing to do with tools. Being into the job. Peers who back your ideas. Useful feedback. The Developer Thriving study (1,282 developers, 12+ industries) landed in the same place.
I say this to every engineering leader I work with: we aren't computers. You can't throw more megahertz at people and expect linear improvements. Life happens. Context switching has real costs. Happy developers are productive developers—most of the time. If your measurement program only counts commits, you're counting the wrong thing.
One more, because it bites. DORA's 2024 report found that unstable priorities tank productivity and spike burnout, even when leadership and docs are good. If your metrics program reads like a stack ranking—congratulations, you just added another unstable priority.
Measure → Understand → Correlate
This is the framework I use. Most teams do the first step and stop.
1. Measure: establish your baseline. You need to know where you are. Pick metrics that matter to your business and track them consistently. This is the easy part—most teams stop here. A spreadsheet is fine for the first quarter. Seriously.
2. Understand: ask why it's that way. Your deployment takes two hours. Okay, why? Are you doing gradual rollouts to 1% of traffic at a time? Do you have 100 million users where an outage would be catastrophic? Is there a manual approval gate because a regulator said so? Maybe two hours isn't a problem. Maybe two hours is exactly what smart looks like for your situation. I've seen teams obsess over a "slow" 30-minute CI pipeline, only to discover those 30 minutes included the test coverage that had prevented multiple production incidents. Cutting them would've been optimizing in the wrong direction.
3. Correlate: connect it to outcomes, and to humans. Fast deployments are great—unless they're fast because you removed the safety checks. High commit frequency looks impressive, until you realize it's because someone's force-pushing to fix their own mistakes. When McKinsey put out "Yes, you can measure software developer productivity" in 2023, Kent Beck and Gergely Orosz took it apart, and their point is the one to keep: most of these metrics measure effort and output. What you want is outcome and impact.
When you're staring at a number, ask:
- Why is this metric what it is? Not "is it good or bad," but what adds up to this number?
- What would "better" actually mean for our business? Faster isn't always better. Sometimes slower and safer wins.
- What are we optimizing for? Speed? Safety? Consistency? Developer experience? You can't have all of them at once.
- Does this connect to something we care about, or is it just easy to measure?
- Have we asked the humans involved?
Your developers know where the friction is. Sometimes the best metric is a conversation.
The survey
Surveys are the backbone of this, because a lot of what matters (how much is in your head, whether you can focus, whether you like your job) only exists in people's heads. Google has run its engineering satisfaction survey every quarter since 2018. Here's how I run one that engineers don't roll their eyes at.
Keep it short and keep it the same. 15 to 25 questions on a 5-point agree/disagree scale, the same core questions every time so you can see the trend, then two or three open-ended questions at the end. That's where the real answers are. A few I always include:
- • "I can get a change from my laptop to production without waiting on another person."
- • "When I have a technical question, I can find the answer in under 15 minutes."
- • "I have enough uninterrupted time to do focused work."
- • "Build and test feedback is fast enough that I stay on task."
- • "It's easy to understand how our services fit together."
- • Open: "What wasted the most of your time this month?"
Every 8 to 12 weeks. Quarterly works for most orgs. Laura Tacho at DX puts it well: by the next survey, people should be able to notice that something changed because of what they said last time. Worried about survey fatigue at scale? Do what Google does—split engineers into three random groups and survey one group per quarter. Everyone answers once a year and you still get a quarterly read.
The response rate is the metric you earn. A typical developer survey at a mid-to-large company gets around 30%. Companies with a real DevEx program get 80% or more, and that's about where execs should start trusting org-wide numbers. Turns out the survey tool has almost nothing to do with it. People answer when they believe something will happen with the answers. So: anonymous, and say so. Only report groups of five or more. Publish last time's results and what you did about them before you send the next one. Have engineering leadership send it, not HR. Two weeks open, one reminder.
Segment it, don't average it. "3.8 out of 5 on build satisfaction, org-wide" tells you nothing. Cut it by team, by tenure (people under six months see friction the veterans have gone blind to), by platform, by role. The gap between groups is your first project.
Small team? Under 20 engineers, skip the survey infrastructure. Ask three questions in every 1:1 and write the answers down. How do you feel about your ability to get work done? What's blocking you? What would make your day better? Tally the themes once a month. You'll have better data than most surveys produce.
What to pull from your systems
Surveys tell you how it feels. Systems tell you what's happening. You want both—the gap between them is where the interesting stuff is. These are the ones that have paid off for me.
| Metric | Dimension | Where it comes from | What bad looks like |
|---|---|---|---|
| Time to first "200 OK" | Cognitive load | Onboarding logs, platform telemetry | A new hire takes 2+ weeks to ship a first change. A new service takes days to get its first deploy. |
| CI duration (p50 and p90) | Feedback loops | CI system | Over 15 minutes. People batch changes and wander off. |
| PR pickup time | Feedback loops | Git host API | Median over a business day from "ready for review" to first review. |
| Time to answer a question | Cognitive load | Support channel timestamps | Questions in #help sit for hours. The same question shows up every week. |
| Meeting-free focus blocks | Flow | Calendar data (aggregate only) | Fewer than two 2-hour blocks per person per day. |
| Unplanned work share | Flow | Issue tracker labels | Over 30% of what got closed wasn't in the plan. |
| Where people give up | All three | Funnel analytics on internal tools and docs | Guides abandoned at step 6 of 10. Doc searches that go nowhere. |
🎉 0 to 200
My favorite one is a nod to HTTP 200 ("OK"), the goal of any successful API call. Your platform, your API, your onboarding—none of it is working unless a developer can go from zero to their first 200 OK quickly, reliably, and with minimal friction. So time it.
But it doesn't stop there. Tracking what doesn't work is arguably just as valuable, if not more. Learning that it took 10 steps to complete a routine task means we can improve. Learning where people gave up because they couldn't get to a 200 is huge. You can learn a lot from what people tell you. You can learn a lot more from watching what they do.
How does this fit with DORA and SPACE? They're not competing. DORA measures delivery. SPACE keeps you honest about measuring more than throughput. What's on this page is the layer underneath both: why your DORA numbers look the way they do. (DX Core 4 packages the same idea into four buckets with a survey index in the middle. DX claims each one-point improvement is worth about 13 minutes per developer per week. Vendor number, take it as one—but it's a fine starting question set.)
Now do something about it
Measuring without acting is worse than not measuring at all. It teaches your engineers that feedback goes into a void.
Week one, before you open a single chart, read every free-text answer and tag it by theme. Feedback is gold, even when it's noisy. Complaints often mask legitimate patterns. Find the root cause, not just the loudest voice.
Week two, pick one thing. Not five. One thing most engineers will notice when it's fixed. At 8x8, "the docs are in seven places" beat every other complaint. Getting everything into one searchable site with a style guide wasn't glamorous. It was the fix everyone felt.
Week three, publish the results and the plan to the whole org—the good, the bad, and "we're not fixing this yet, here's why." Then say what you are fixing, who owns it, and when they'll hear about it again. If you build it, they won't come. You have to tell them. Chat, all-hands, the onboarding site, email. The onus of communication is on the communicator, and in a remote world there's no harm in over-communicating.
Then fix it, and watch the number that should move (CI p90, time to first deploy, questions per week in #help). When the next survey goes out, you want to open with "you told us X, we did Y, here's what changed." That's what gets you the 80% response rate.
Measure everything against two pillars: Vision (are we enabling what we set out to?) and Value (is it actually making life better?). If a metric doesn't help answer one of those, stop collecting it.
Where this goes wrong
I've watched this fail the same few ways.
Measuring individuals. Commits per engineer on a dashboard the VP can sort by name. Don't. Team level and org level only. When a measure becomes a target it stops being a good measure—Goodhart's law, and I've never seen it lose. If people are gaming your metrics, you built the wrong incentive.
Cargo-culting the dashboard. Google's own engineering book starts every measurement with one question: what decision will this change? If the answer is "none," don't measure it.
Surveying and then doing nothing. The annual engagement survey goes out, results land in a deck, nothing changes, and next year the response rate drops from 40% to 25%. Don't send a survey without engineering time already budgeted to act on it. If you can't promise to fix at least one thing, don't ask.
Trusting only the numbers, or only the feelings. "The data says CI is fast" while everyone complains about CI. Google's productivity team checks the logs against what people report, and the other way around. When they disagree, that's the finding. Go talk to people.
And the new one: "We rolled out AI coding tools, so productivity is up." DORA's 2024 data tied AI adoption to small gains in flow and satisfaction—and to lower throughput and noticeably lower delivery stability. Their 2025 report's line is that AI amplifies whatever your org already is. Measure stability next to adoption or you'll throw a party for a regression.
If you're starting Monday
Talk to ten engineers across teams and tenure this week. Shadow one new hire from laptop to first merged change and count the waits. Pull the free data—CI times, PR pickup, deploy counts—into a spreadsheet.
By the end of the month, run the first survey. By day 45, publish the results and commit to one fix with an owner and a date. By day 90, ship it, send the second survey with a "you told us, we did" preamble, and compare response rates. If they went up, it's working. If not, you skipped a step.
Measuring is easy. Understanding is hard. That's why Google, Microsoft, and GitHub have whole teams for this. Don't look at your metrics and wonder if they're "good enough." Look at them and ask what they're telling you about how work gets done here. The goal was never to stack rank anybody. It's to find the papercuts, reduce friction, and help people do their best work.
If you're wrestling with this at your company—tracking a bunch of stuff and not sure you're learning anything—drop me a note. I'd love to compare notes.
-Matt
Sources
- Noda, Storey, Forsgren, Greiler. "DevEx: What Actually Drives Productivity." ACM Queue, 2023.
- Forsgren, Kalliamvakou, Noda, Greiler, Houck, Storey. "DevEx in Action: A Study of Its Tangible Impacts." ACM Queue, 2024.
- GitHub. "The Good Day Project." Octoverse Spotlight, 2021.
- Murphy-Hill, Jaspan, et al. "What Predicts Software Developers' Productivity?" IEEE TSE, 2019.
- D'Angelo et al. "Measuring Developer Experience With a Longitudinal Survey." IEEE Software, 2024.
- Jaspan. "Measuring Engineering Productivity." Software Engineering at Google, 2020.
- Hicks, Lee, Ramsey. "Developer Thriving." Developer Success Lab, 2023.
- Atlassian. State of Developer Experience Report 2024 and 2025.
- Stack Overflow. Developer Survey 2024 and 2025.
- DORA. Accelerate State of DevOps Report 2024 and State of AI-assisted Software Development 2025.
- DX. "Measuring Developer Productivity with the DX Core 4." 2024. Developer survey participation rates.
- Tacho. "A Deep Dive into Developer Experience Surveys."
- Beck and Orosz. "Measuring Developer Productivity? A Response to McKinsey." 2023.
- Matt Gardner. "Everyone's Measuring Developer Productivity. Almost No One Knows Why." and "Building a Platform That Doesn't Suck."
Tracking Metrics But Not Learning Anything?
I've been thinking about this stuff for 15+ years, from Apple to 8x8, and I'm still learning from what other people are seeing. Tell me what you're tracking and where it's not adding up.
Get in Touch