工程仪表盘可能在骗你:OpenVector 如何用指标关联揭示真实工程效能
Your Engineering Dashboard May Be Lying to You
资深工程负责人指出,自动化覆盖率从 60% 升至 85%、代码覆盖率达 90% 等单项指标改善,可能同时伴随缺陷泄漏率从 8% 升至 14%、生产事故增加,指标孤立查看会产生误导。AI 编码助手让代码产出翻倍,也不等于工程价值翻倍。为此作者开源了 OpenVector,从交付、质量、自动化、代码健康、容量与 AI 影响六个维度关联指标,让工程信号之间的关系可见,而非再堆一个图表仪表盘。
Your Engineering Dashboard May Be Lying to You
Velocity is up.
Automation coverage is up.
Code coverage is up.
Defects are down.
Everything looks great.
But are we actually building better software?
That question has been on my mind for a long time.
After working across Engineering, Quality Engineering, Delivery, DevOps, cloud transformation and engineering leadership, I've noticed a recurring problem:
We measure engineering performance using metrics that are individually useful but collectively misleading.
A metric can improve while the engineering outcome gets worse.
When 85% automation coverage isn't necessarily better
Imagine an engineering organization moves from:
60% automation coverage → 85%
The dashboard shows a clear improvement.
But at the same time:
- Defect leakage increases
- Production incidents increase
- Test maintenance effort increases
- Critical user journeys remain poorly covered
Did automation actually improve?
The answer isn't obvious anymore.
The problem isn't the automation metric.
The problem is looking at it in isolation.
What about 90% code coverage?
Code coverage is another good example.
Suppose a team reaches:
90% code coverage
That sounds impressive.
But coverage doesn't tell us:
- Whether the right scenarios are being tested
- Whether assertions are meaningful
- Whether critical business paths are protected
- Whether the code is maintainable
- Whether defects are escaping into production
Two codebases can both report 90% coverage and have very different levels of quality.
Again, the metric isn't necessarily wrong.
Our interpretation of the metric may be incomplete.
And then AI changes the equation
AI makes this even more interesting.
Imagine an engineering team using AI coding assistants and suddenly producing twice as much code.
That's a significant productivity signal.
But did engineering value increase by 2×?
Not necessarily.
We also need to understand:
- Defect trends
- Code quality
- Review effort
- Technical debt
- Test effectiveness
- Production reliability
- Customer impact
AI can increase the amount of software produced.
The harder question is whether it increases the value of software delivered.
The problem with engineering dashboards
Most engineering organizations have no shortage of metrics.
They measure:
Delivery
- Sprint velocity
- Planned vs. completed work
- Cycle time
- Capacity utilization
Quality
- Defects
- Defect leakage
- Severity
- Defect aging
Automation
- Automated tests
- Automation coverage
- Execution frequency
- Test pass rates
Code
- Code coverage
- Code quality
- Technical debt
- Maintainability
Engineering productivity
- Pull requests
- Review time
- Deployment frequency
- Lead time
AI
- AI-assisted coding
- Code generation
- AI test generation
- Developer adoption
All of these can be useful.
But here's the problem:
They are usually viewed as separate dashboards.
That's where the real opportunity is.
The signal is in the relationship between metrics
Consider a simple example.
| Metric | Previous | Current |
|---|---|---|
| Velocity | 40 SP | 55 SP |
| Automation | 60% | 85% |
| Code Coverage | 72% | 90% |
| Defect Leakage | 8% | 14% |
If we look at the first three metrics independently, the story looks fantastic.
But once we connect them with defect leakage, the story becomes much more interesting.
Something changed.
And that's the question an engineering leader should investigate:
Why?
Maybe the organization is optimizing for delivery metrics.
Maybe automation is focused on the wrong areas.
Maybe tests are providing coverage without meaningful validation.
Maybe AI-generated code increased output faster than review and quality processes could adapt.
Maybe the organization is simply measuring the wrong things.
The dashboard doesn't give us the answer.
But it should help us ask the right question.
From measuring metrics to understanding engineering systems
This is the thinking behind OpenVector — Software Engineering Metrics Unleashed, an open-source project I'm building.
The idea is simple:
Don't just collect more metrics. Connect the metrics you already have.
OpenVector looks at engineering through multiple dimensions:
📈 Delivery
Sprint velocity, capacity, planned vs. completed work
🧪 Quality
Defects, leakage, severity, aging and sources
🤖 Automation
Automation coverage and execution trends
💻 Code Health
Code coverage and code quality
👥 Capacity
Engineering capacity and productivity signals
🧠 AI Impact
Understanding AI adoption alongside engineering outcomes
The objective isn't to create another dashboard full of charts.
It's to make relationships between engineering signals visible.
The metric isn't the outcome
This is perhaps the most important distinction.
Velocity is not value.
Automation coverage is not quality.
Code coverage is not test effectiveness.
AI-generated code is not engineering productivity.
A declining defect count is not automatically improving quality.
These metrics are signals.
The engineering outcome emerges from how those signals interact.
That's why I believe the next generation of engineering intelligence needs to move from:
Metric → Dashboard
to:
Metrics → Relationships → Insights → Decisions
What should engineering leaders measure?
I don't think the answer is simply "more."
In fact, I think we should ask a different question:
Which relationships between metrics help us make better engineering decisions?
For example:
Velocity + Defect Leakage
Can increased delivery speed be achieved without sacrificing quality?
Automation + Defect Leakage
Are we automating the right things?
Code Coverage + Defects
Is increased coverage translating into fewer escaped defects?
AI Adoption + Code Quality
Is AI-assisted development improving engineering outcomes or simply increasing output?
Capacity + Delivery
Are teams becoming more efficient, or are they simply working harder?
Those relationships are much more interesting than any individual number.
More metrics ≠ more insight
Engineering organizations don't necessarily need another hundred metrics.
They need better ways to understand the metrics they already collect.
More metrics ≠ more insight.
Better-connected metrics = better decisions.
That's the problem I'm exploring with OpenVector.
If you're an engineering or technology leader, I'd be interested in your perspective:
Which engineering metric do you trust most when making a serious decision?
And perhaps more importantly:
Which metric does your organization measure simply because it's easy to measure?
EngineeringMetrics #SoftwareEngineering #QualityEngineering #AI
来源:Google AI:DEV 作者专属(RSS) · dev.to