The Lean AI Framework™: Stop Slop, Accelerate Value
Note: though this article is about AI, every word was written by and for humans for maximum reading enjoyment.
The top question boards and executives are asking engineering leaders in 2026 is: “What’s the impact of AI?” They asked in 2025 too, but the truthful answer last year was: “Meh, it only costs $20 a month though.”
A lot has changed since the release of Opus 4.5, Gemini 3, and GPT-5.2 at the end of 2025. AI works now, and it’s expensive. Gartner found monthly bills of $2,000-$5,000 for heavy users. minware’s engineers have had their fair share of $500 days.
Despite the massive uptick in AI use, it is not all sunshine and rainbows. DORA’s report on the State of AI-Assisted Software Development highlights how AI magnifies both good and bad practices. In July 2026, The Pragmatic Engineer warned of a massive increase in code review load.
In the widespread euphoria of AI adoption, everyone from individual engineers up to executives has been pushing for higher code output. Influencers on LinkedIn brag about their token spending, while AI vendors happily encourage running agents on as many parallel tasks as possible.
AI measurement frameworks from other engineering analytics platforms reinforce this output-first view, putting metrics like merged pull request count front and center.
All the hype has left everyone disillusioned, with engineers drowning in slop while business leaders see ever-increasing PR counts with no corresponding increase in business value.
This lack of productivity gain should not come as a surprise. It’s a well-known principle of lean methodology that high utilization hurts productivity in knowledge work. The rapid runup in agent activity has ballooned human utilization with code reviews, agent intervention tasks, and context switching.
The software industry needs a new way to think about AI productivity.
Introducing the Lean AI Framework
minware’s Lean AI Framework™ offers a fresh look at AI use from long-established first principles of lean methodology. The framework emphasizes increasing value delivery by reducing slop, which we define as work or output that consumes tokens or human time without contributing to delivered value. In lean parlance, slop is equivalent to “waste.”
The Lean AI Framework’s approach to AI productivity runs contrary to the prevailing industry narrative. Instead of pushing for more agent sessions working on tasks in parallel that drive up human context switching, its model scenario is AI completing tasks that have a human in the loop one at a time per person (single-piece flow) as quickly as possible with high quality, low rework, and minimal human intervention. This cuts down on slop while tightening the loop of customer value delivery and learning.
The Lean AI Framework also goes one level deeper than other frameworks, which only measure outcomes. Other frameworks leave teams on their own to figure out how to make AI more effective. In the Lean AI Framework, best practice metrics that give individual teams a roadmap for improving their AI productivity are first-class members rather than an afterthought.
Finally, metrics in the Lean AI Framework are designed to be customized for the unique way each team operates and creates value for customers. This is essential for distinguishing between slop and value – something off-the-shelf metrics like pull request counts fail to do.
Customization previously required an in-house analytics solution. minware’s patent-pending minQL formula language and hypercube data model allow customer success to tailor metrics to each team’s operating model during onboarding, and let teams edit the metric logic themselves in the application. This offers the best of both worlds: the flexibility of an in-house solution without having to build it yourself.
These three differentiators – fast flow, best practices, and customization – make minware’s Lean AI Framework uniquely able to help engineering leaders, managers, and individual contributors maximize the impact of AI on value delivery while putting a stop to slop.
Pillars of the Lean AI Framework
minware’s Lean AI Framework consists of three pillars, organized by the highest level of the organization that consumes the reports and how the metrics build upon one another.
| Pillar | Highest Reporting Level | Focus |
|---|---|---|
| Business Outcomes | Board, CEO, CFO | Value, Cost |
| Operational Outcomes | CTO, Engineering Leaders | Quality, Efficiency, Predictability |
| Best Practices | Engineering Teams | Team Workflow Efficiency, Planning and Prompting, Session Automation, Culture |
At the top, business stakeholders like the board, CEO, and CFO primarily care about AI’s impact on business outcomes: value, cost, and return on investment (ROI).
To drive increased AI ROI, engineering leaders have to look one level lower at operational outcomes that are leading indicators of value delivery, like bug creation, cycle time, and sprint completion.
Other AI frameworks stop here.
minware’s Lean AI Framework goes deeper by looking at best practices that help engineering teams understand the root cause of deficient operational metrics and act as leading indicators. For example, looking at work in progress (WIP) by ticket status can identify bottlenecks that drive up cycle time, and analyzing human session interventions can uncover opportunities to make agents faster and more reliable.
The sections that follow describe how software organizations can use the metrics in each pillar to drive AI’s impact on value delivery.
Pillar 1: Business Outcomes
The first pillar of minware’s Lean AI Framework covers business outcomes. It answers the number one question that boards, CEOs, and CFOs are asking engineering leaders: what is the value created by AI, and what does it cost?
At the end of the day, if AI ROI is strong, business leaders are not likely to care about lower-level metrics like cycle time and sprint completion. Conversely, if AI ROI is negative, then no amount of stellar operational metrics will make up for it. Business outcomes are typically shown in reports at the board, CEO, and CFO level.
This pillar consists of the following metrics, which are explained in the sections that follow.
| Goal | Metric |
|---|---|
| Value | Roadmap Value Delivery |
| On-Time Delivery Rate | |
| Cost | Roadmap Delivery Cost |
| Overhead Cost |
Analysis methodology for these metrics is covered in the AI Impact Analysis Methodology section.
Roadmap Value Delivery
Every software organization that plans engineering work estimates its expected value, whether or not those estimates are explicit. The way each organization represents expected value may differ, ranging from a simple ranked list to sophisticated data science models. Even if there is no defined process, a value estimate implicitly exists in the mind of the person who sets priorities, and the CEO must assess value to judge the engineering leader’s performance.
The roadmap value delivery metric reflects the expected value of completed work items. Customers may configure this metric in minware to connect to any field from their project management system where they set explicit value estimates during the roadmap planning process. It may also be loaded from spreadsheet custom uploads for teams that record it outside of their formal project management system.
If no explicit value field is configured, minware defines roadmap value delivery as story points completed on tickets that are not bugs and have a parent epic/project ticket, defaulting to 1 per ticket if no estimate is set. When roadmap tasks are kept in prioritized order based on expected ROI (value / cost), people always work on the highest priority tasks, and the relative ROI of the best opportunities stays stable over time, cost estimates can be a reasonable proxy for value. That being said, AI may break these assumptions by altering per-item ROI, causing people to pick up lower-priority work further down the roadmap, or changing the relationship between story points and cost for tasks that become easier. So, hooking up true value estimates is preferable.
Also, many customers configure which tickets count toward roadmap value rather than relying on epic membership. Those doing cost capitalization may use an existing “capitalizable” field. Others may use a custom formula based on issue types, labels or other fields to align with how each team defines value. This is particularly important for operationally focused teams that do not often work on projects.
However roadmap value delivery is defined, it’s important to keep in mind that it is an expected value, and that the actual value may diverge. The framework uses expected value because it’s more readily available. Actual value is often not feasible for organizations to accurately measure over a short time horizon or attribute to individual roadmap items. So, in addition to looking at AI’s impact on expected value with roadmap value delivery, we recommend reconciling those expectations to actual value where possible, such as with A/B test experiments that measure business KPI impact. If actual value is available in a spreadsheet or custom ticket field, minware can display it side-by-side or instead of the expected value in roadmap value delivery.
The roadmap value delivery metric is critically important because it forms the basis for assessing AI impact. Aligning the whole organization around roadmap value delivery as a north star is crucial for minimizing slop because it sends the message to every engineer, manager, and leader that whatever they do will be judged by the value they create for customers, not their amount of output.
On-Time Delivery Rate
Roadmap value delivery is necessary for understanding AI impact, but it is not sufficient. For most organizations, delivering roadmap projects on time is an essential component of value. Late projects can cause missed opportunities and negatively impact the rest of the organization.
The on-time delivery rate metric captures the loss associated with missed deadlines. It helps leaders assess whether AI is having a positive or negative impact on the engineering team’s ability to deliver on its commitments. It is a more precise manifestation of the simpler Say-Do Ratio.
Missed deadlines may have different levels of impact depending on the nature of the project and organization. For some, being late by even one day could negate the entire value of the project, such as missing an external deadline for submitting bids. For other projects without external dependencies, value may degrade more slowly, such as by generally missing out on opportunities to capture more market share in the current quarter.
minware’s default definition of on-time delivery rate is the estimated roadmap item duration (time from when a roadmap item is marked in progress until the original due date set at that time) divided by the actual roadmap item duration (time from in progress until it is marked done), capped at 100% per roadmap item. By default, minware considers each epic a roadmap item, but this can be configured to use higher-level tickets instead. This heuristic discounts project value gradually as delivery slips, and doesn’t give any extra credit for early delivery.
During the onboarding process, minware’s customer success team can calibrate this metric with customers to reflect the reality of their organization, such as by using a sharper drop-off or immediately counting missed deadlines as zero. The metric may also be configured to include incomplete projects that are currently overdue in the calculation, as well as those that are entirely canceled or abandoned.
Roadmap Delivery Cost
The next metric in the business outcomes pillar is roadmap delivery cost. This metric includes both engineering personnel costs and AI token costs. It is important to count both of these because ideally AI should replace human effort with less expensive tokens, driving down the total cost of each roadmap item.
To assess ROI at a project level, human and AI roadmap delivery costs should be attributed to individual epics or parent tickets because different projects may have vastly different amounts of AI usage.
minware computes cost per roadmap item by combining personnel costs from its patent-pending time model with token costs from AI tools like Claude Code, Cursor, Codex, and Copilot. minware attributes human and AI costs to each commit, then rolls up the costs to pull requests, tickets, and epics via its patent-pending hypercube data model. minware breaks down human vs. AI cost in its reports so that it’s easy to see how heavily each project used AI.
For many AI tool plans, billed costs differ from per-token costs. AI tool APIs and OpenTelemetry data available to minware generally report what the token cost would be on a consumption-based plan with no discount.
If you have significant discounts in place, we recommend configuring the roadmap delivery cost metric to multiply reported token costs by your total bill divided by total reported spending in minware to reflect your average discount.
It’s also important to keep in mind that certain vendors like Anthropic have significantly different pricing structures for enterprise plans. If you expect to cross a threshold that changes your pricing tier in the near future, such as a seat count that moves you to an enterprise plan, you may want to report current impact metrics based on the future enterprise cost.
Overhead Cost
No engineering team spends 100% of their time on value-adding work. In addition to roadmap commitments, engineering teams are responsible for fixing bugs and maintaining existing software, plus hiring, career development, and other corporate activities.
The overhead cost metric captures the personnel costs and AI token spend that goes toward keeping the lights on. For organizations that do cost capitalization, the overhead cost is the amount of non-capitalizable human effort and token spend.
By default, minware computes the overhead cost by combining personnel costs from its patent-pending time model with AI token costs, and then looking at costs that are either not traceable to any ticket (i.e., if a person has no assigned, in-progress ticket), or associated with a ticket that has been excluded from consideration in the roadmap value delivery metric (by default, a ticket that is a bug or does not have a parent epic/project ticket). minware also breaks down overhead cost by human vs. AI costs to show how much AI is helping to automate non-value-adding work. It also shows overhead cost as a percentage of total cost to display the non-capitalizable work rate, though this rate shouldn’t be regressed against token spend because token spend is part of the denominator, so changes to it will change the rate even when overhead is unchanged.
During onboarding, organizations typically work with customer success to customize the overhead cost metric, such as by deriving it from existing “capitalizable” fields in Jira, or creating formulas based on issue types or other fields. (The overhead vs. value determination must be kept in sync and complementary between roadmap value delivery and overhead cost.) Customers often define different categories of overhead as well to differentiate between major sources like bugs, maintenance, and technical debt.
Overhead cost is an often-overlooked but potentially major component of AI impact. AI is very good at speeding up repetitive maintenance work and burning through tech debt tasks like version upgrades. It’s important to look at whether overhead is decreasing overall, and how much of the overhead burden is shifting from engineers to AI.
Pillar 2: Operational Outcomes
The next pillar of minware’s Lean AI Framework focuses on operational outcomes. These outcomes reflect the overall health of the software development lifecycle (SDLC).
Operational outcome metrics do not directly measure value, but serve as leading indicators to help CTOs and engineering leadership understand what is driving value and where to improve. They also act as guardrails to ensure that AI business impact isn’t coming at the expense of a decline in quality.
Operational outcome metrics also help engineering leaders assess whether individual teams are performing well without having to dig into best practices. Individual teams and engineering managers generally report these metrics to their engineering leaders alongside business outcome metrics. Engineering leaders often do not report operational outcome metrics to the CEO or board unless there is a problem with business outcomes, or if business leaders want to verify that proper guardrails are in place.
The operational outcomes pillar consists of the following metrics, which are explained in the sections that follow:
| Goal | Metric |
|---|---|
| Efficiency | Pull Request (PR) Lead Time |
| Ticket Cycle Time | |
| Quality | Bug Rate |
| Change Failure Rate | |
| Bug Lead Time | |
| Predictability | Sprint Completion |
For each of these metrics, we recommend assessing AI impact with the methodology described in the AI Impact Analysis Methodology section.
Bug rate and change failure rate come with additional caveats: a higher level of quality problems could cause more token spending on remediation. Bug rate can also increase without a true decrease in quality if AI is used to find more bugs. As a result, we recommend carefully looking at whether total spending on bugs changed and how many new bugs were found with AI when interpreting these correlations.
Pull Request (PR) Lead Time
In lean methodology, lead time is one of the most important metrics. It measures the total time from concept to value delivery. A short average lead time means that the value delivery process is free from significant delays caused by quality issues or other workflow bottlenecks.
In software development, the inner-most value delivery loop is writing and deploying code. Pull request (PR) lead time is a narrower measure, commonly defined as the duration between the first commit and deployment for each code change that goes through the pull request process. It includes time for code review, integration with other changes, and final deployment to production.
minware defines PR lead time following the common definition as the time from first commit to deployment. Deployment time defaults to the time code is merged into a main branch, but can also be configured to measure deployment based on Git tags or CI/CD deployment pipeline runs.
minware’s reports break down lead time by stage so that you can see if AI is having a disproportionate impact on different parts of the pull request process. AI tools are notorious for overwhelming engineers with slop code review. This breakdown tells you the impact of AI on code review in your organization.
PR lead time is critically important for assessing AI impact because it captures many ways AI can accelerate development or drag it down with slop. There are a lot of best practices that go into effective AI use such as planning, tool use, and test automation, but almost all of them will show up in PR lead time metrics.
Ticket Cycle Time
The weakness of using PR lead time to assess AI impact is that it only reflects the inner-most loop of the software development lifecycle. A common theme we hear from customers is that AI is helping them write code, but they’d like to see it accelerate other tasks like requirements gathering and user acceptance testing.
The ticket cycle time metric provides a broader gauge of how long each work item takes to go from start to finish. minware defines ticket cycle time as the time between when a ticket moves to an in-progress status and when it is marked done.
Customers often configure ticket cycle time by filtering or segmenting the type of tickets to which it applies (e.g., ignoring subtasks or separately looking at cycle times for stories and parent epics), or adjusting the definitions of in-progress and done.
minware also breaks down ticket cycle time by status so that you can see exactly which workflow steps are delaying value delivery.
PR lead time tells you whether engineers are effectively using AI to write code. But, when gains in code output fail to move roadmap value delivery, ticket cycle time helps you figure out why.
Bug Rate
Another core tenet of lean methodology is driving down defects and pushing them as close as possible to their source. In software, a defect that makes it into production and results in a bug requires much more rework than an issue caught during testing.
Also, it’s possible for both PR lead time and ticket cycle time to get better while productivity suffers if using AI results in more bugs.
The bug rate metric is an essential guardrail for AI impact assessment that captures the positive or negative effect of quality on the software development process.
minware measures bug rate by dividing the total number of bugs created by the number of code changes. By default, minware counts all new bug tickets and divides that by the number of merged pull requests.
Some organizations customize the bug rate metric during onboarding to distinguish between bug tickets created prior to release from those in production, making it effectively an escaped defect rate. Others choose to use a different denominator like story points completed to minimize the risk of a decline in quality being obscured by increased PR volume, whether from AI tools or from splitting PRs to reduce PR complexity.
While bug rate doesn’t cover issues that are caught by testing prior to becoming a bug, it captures quality problems that have the biggest impact on value delivery and may not show up in other efficiency metrics.
Change Failure Rate
The bug rate tells you how many defects go into production, but it doesn’t account for their severity. Severe bugs that cause production incidents have a greater impact on customers and are much more disruptive to the software development process.
The change failure rate (CFR) metric captures the differential impact of severe bugs by measuring the number of change failures (high-priority incidents) divided by the total number of changes.
By default, minware defines change failure rate by dividing the number of highest-priority bugs by the number of pull requests merged to a main branch.
Customers often configure CFR during the onboarding process to specify which priorities, labels, or other fields indicate a change failure in the project management system, or pull incident counts from a helpdesk system like ServiceNow.
Customers may also configure the change count denominator to be production deployments indicated by Git tags or pipeline runs from an integration with a CI/CD system like GitHub Actions. Denominators like story points completed may also be used to reduce the risk of higher PR volume from AI tools or PR splitting distorting this metric.
Change failure rate is an important guardrail metric for assessing AI impact beyond bug rate. It tells you whether use of AI tools is leading to more or fewer severe problems that escape the QA process.
Bug Lead Time
Bug rate and change failure rate alone overlook a critical component of quality: how long bugs take to fix. The longer a bug sits around, the greater its impact on customers and the more work it is to remediate. Latent bugs with many intervening changes tend to be harder to isolate and unwind, causing greater disruption to value delivery.
Bug lead time measures the time from when a bug was reported to when it was fixed. By default, minware calculates bug lead time as the time from ticket creation to ticket resolution for all bug tickets.
Customers may configure bug lead time to start earlier at the time of introduction if that information is recorded in a custom field, or end later at the time when the fix is deployed if that is later than ticket resolution. It can also be helpful to break down bug lead time by stage to identify bottlenecks in the workflow of discovery, reproduction, assignment, implementation, review, and deployment.
The nice thing about bug lead time is it’s one area where AI has the opportunity to make a big impact. Linear built an agent to investigate and automatically fix bugs. Our personal experience at minware is that Claude Code is great at isolating the root cause of complex bugs that would have taken an engineer hours to find before AI.
Sprint Completion
It is possible for teams to have great cycle/lead times, excellent quality, and still fail at value delivery. How? Predictability problems like interruptions or planning/estimation misses that result in scope creep can cause teams to fall short of their original goals.
A common refrain we hear from our customers is: “we can’t get anything done because of all the other stuff that gets in the way.”
The agile scrum process is designed to combat this exact problem. It involves measuring completed work against the original commitment in fixed iterations commonly known as sprints (though some teams and project management tools refer to them as “cycles” or “iterations”).
Sprint completion measures the amount of completed work divided by the original commitment. It reflects the team’s ability to stay focused and make progress on planned value-adding activities.
Sprint completion is an earlier leading indicator of on-time project delivery. Both target predictability, but problems with sprint completion will show up sooner and allow teams to take corrective action before projects are late.
By default, minware defines sprint completion as the total number of story points completed each sprint divided by the story points in the original commitment, capped at 100% per sprint. This is a less strict definition because it does not penalize late additions to the sprint.
Note that the sprint completion metric does not address undercommitment. If teams commit to far less work than they are capable of completing, then their sprint completion metrics will look good. We rarely see this issue in practice, and the far more common problem is overcommitment. That being said, unusually consistent 100% sprint completion warrants deeper investigation into specific completed work items to assess whether teams are undercommitting or gaming metrics in other ways like shifting the definition of done.
Customers often configure this metric in minware to reflect different choices in their own process, such as how to count tickets that are missing estimates, whether to give credit toward completion for tickets added to the sprint after it started, and the definition of done. Some teams also use a custom field or label to distinguish between added tickets that are interruptions vs. deliberate scope changes and only discount the interruptions from counting toward sprint completion.
minware’s reports also break down tickets that did not complete by status and by the number of sprints they have rolled over to identify specific bottlenecks that interfere with predictability.
Teams not using sprints and following a Kanban process may replace this metric with ticket or story point completion vs. a fixed weekly commitment, which is an analogous predictability measurement for non-scrum processes.
Sprint completion may not tell you AI’s impact on throughput if velocity gains result in higher commitments, but it answers another important question: is AI causing teams to be more or less focused and predictable? With AI’s role as an amplifier, this one can go either way. Some teams may automate interruption-handling and plan more thoroughly with AI, while others may get sidetracked and pursue even more distractions because they think AI will make it easier.
AI Impact Analysis Methodology
This section covers the framework’s recommended analysis methodology for assessing AI impact on business and operational outcome metrics.
Conceptually, net engineering return on investment follows this formula:
1 + Net ROI = (Roadmap Value Delivery × On-Time Delivery Rate) / (Roadmap Delivery Cost + Overhead Cost)
This formula can be computed at the roadmap item level with all values other than overhead cost, which is generally amortized in proportion to each item’s roadmap delivery cost, using overhead from the same team and time period.
When assessing AI impact and incremental ROI driven by AI, we recommend separately running linear regressions of roadmap value delivery, on-time delivery rate, and overhead cost against AI token spend, which is the variable component of roadmap delivery cost for consumption-based plans.
Looking at each metric separately better shows the correlation of token spend with each one and is easier to reason about than the combined net ROI formula, especially since value delivery may be represented in abstract units rather than dollars. (Abstract units make net ROI undefined and must be translated to dollars to assess final profitability.) The combined formula is important to understand if questions arise, but AI impact reports typically just show the relationship between token spend and the other three metrics independently.
A linear regression tells you the incremental change in value associated with each additional dollar’s worth of token usage (though actual costs may not kick in until you reach overage thresholds with certain plans like Claude Code Pro or Max). minware’s reports show regressions by default grouping by team and time period and normalizing each data point to be per-person-day by dividing token spend on the X axis and count or dollar metrics on the Y axis by active contributor work days computed via minware’s time model.
minware’s reports further show scatter plots and list the individual data points so you can zoom in on those with the highest spending to assess whether they are creating value or slop.
Correlation vs. Causation and Controlling for Confounding Variables
Other frameworks utilize AI vs. non-AI cohort comparisons when analyzing AI impact, which we don’t recommend for two reasons. First, non-AI control groups are becoming less meaningful or disappearing entirely with the widespread adoption of AI.
Second, to the extent non-AI cohorts remain, they are increasingly prone to selection bias. If an engineer chooses not to use AI, it’s probably because it’s unlikely to help, such as in a legacy repository with poor test coverage. AI use may correlate positively with business outcomes in cohort comparisons when the true causal link is with code quality rather than the incremental effect of AI.
However, causality and selection bias are a concern with regressions as well. Covariates like code base age, code quality, and engineer experience level may influence decisions by the engineer about how much to use AI and the resulting business outcomes.
This can lead to correlations that are not driven by the incremental impact of AI. For example, an engineer may spend more on tokens working in a repository without tests that also has worse delivery outcomes, which could make the impact of AI look negative even if it was neutral or helpful.
For more in-depth analysis, you can configure minware’s reports to add potentially confounding variables and export the full data set to CSV for further investigation in a statistical analysis tool like R.
Additionally, we recommend looking at the value of each business outcome metric over time for the same team to see its trend as AI maturity increases and control for differences between teams (though differences within teams like different projects or personnel changes remain). minware shows historical trend charts alongside linear regressions in its reports to support this type of analysis. Custom groupings like repository or roadmap area are also configurable to control for different effects.
Sample Sizes and Data Interpretation
When running linear regressions, it’s essential to have enough data points to distinguish a real slope from noise. 20 to 30 data points are recommended. For short time periods, within-team analyses, or organizations with a small number of teams, it’s easy to end up with insufficient data. To address this, we recommend extending the time range rather than narrowing each time period, and comparing data from multiple teams instead of looking at each team individually if necessary.
You should also drill down into outliers in the regression to understand if they reflect the rest of the population. For example, someone with massive token spending on a one-off experimental project may be worth excluding with filters to prevent the data point from influencing the result, though exclusions should be noted in reports.
It’s also important to note that some of the business outcome metrics and operational metrics are bounded. On-time delivery rate, sprint completion, and change failure rate are capped at 100%. A straight-line fit can predict impossible values for these metrics and understate uncertainty near the bounds, particularly if there are a lot of 100% points. Bug rate is non-negative but unbounded above, so a straight-line fit can predict negative values. For a formal analysis, you can export the data and use a different model suited to proportions or counts in an external tool.
Finally, minware’s reports show an R² with each linear regression that tells you how much of the variation is explained by token spend. A low R² value means that the relationship is weak even if the slope is positive, so you should report it along with the slope to avoid overstating the results. (Though R² is unreliable below 20 data points and can indicate a strong fit while predicting impossible values for capped metrics.)
Analyzing High-Cost Outliers
In addition to looking at overall ROI trends, investigating particular outliers can shed light on specific work items with exceptionally high cost / value ratios that are most likely to contain major root causes of slop.
minware’s reports surface lists of individual epics and tickets with the highest roadmap delivery cost relative to their roadmap value delivery metric. This gives you a short list for further investigation to assess whether the spending was justified or wasteful.
minware’s reports also show zero-value items with the highest token and time costs so that you can identify the biggest sources of spending on overhead.
Pillar 3: Best Practices
The next pillar of minware’s Lean AI Framework shifts focus from the AI’s effect on outcomes to the best practices teams can apply to improve those outcomes.
These best practice metrics give individual contributors and engineering managers actionable day-to-day guidance for optimizing AI impact on value delivery and driving down slop.
The best practices covered in this section run counter to how many people in the software industry talk about AI productivity. Instead of running parallel agent sessions on as many tasks as possible and driving up human context switching, these best practices adhere to lean methodology. They guide you toward improving your AI workflow by speeding up agent sessions and increasing quality to reduce rework. The ultimate goal is to enable focused effort on a single task at a time with AI.
The best practices pillar consists of the following metrics, which are covered in more detail in the sections that follow. Unlike higher-level metrics, best-practice metrics target specific sources of slop/waste, which are listed here as well.
| Goal | Metric | Main Slop Mitigated |
|---|---|---|
| Team Workflow Efficiency | Ticket Work-in-Progress (WIP) | Human context switching |
| Planning and Prompting | Ticket Size | Wasted tokens on underspecified functionality |
| Pull Request (PR) Complexity | Human review time on avoidable complexity | |
| Pre-Merge Churn Rate | Wasted tokens on discarded code | |
| Post-Merge Churn Rate | Wasted tokens on short-lived code | |
| Session Automation | Human Commit Rate | Human time doing tasks that could be automated |
| Agent Session Length | Human time waiting for agents | |
| Human Intervention Rate | Human time babysitting agents | |
| Cloud Agent Use Rate | Human time managing sessions | |
| Cloud Agent Success Rate | Human time fixing problems | |
| Automation Runtime | Human time waiting on agents that are waiting for tools | |
| Culture | AI Rules of Engagement | Human review time on avoidable quality problems |
Note that the AI rules of engagement metric is the only one assessed outside of minware.
When analyzing these metrics, we recommend looking at their trend over time individually for each team. Within-team comparison is preferable to cross-team comparison because team differences may reflect varying metric definitions, workflows, and responsibilities. For example, comparing a team that maintains a legacy application to one working on a greenfield project wouldn’t be apples-to-apples.
It can also be counter-productive to segment out non-AI tickets or pull requests when looking at these metrics because improving them for tasks that are currently non-AI can increase AI’s ability to take on similar tasks in the future.
Best-practice metrics are diagnostics for improving impact rather than impact metrics on their own, so correlating them with token spend using a linear regression is not necessary or recommended. It would also be difficult to infer the direction of causation. Improving these best practices is likely to cause a decrease in token spending by decreasing waste, but it may increase spending on value-adding work. Using AI well or poorly can also increase or decrease these metrics with causality in the other direction.
Note on Data and Metric Availability
The ticket WIP and ticket size metrics in this pillar require a ticket data integration. PR complexity, pre-merge churn, post-merge churn, and human commit rate rely on a Git integration. AI rules of engagement do not require any data integration.
Aside from human commit rate, the session automation metrics in this pillar and the active agent session time metric used for assessing automation runtime require granular agent session data. Human commit rate will show data, but the results will be incorrectly inflated.
With Claude Code, Codex, and GitHub Copilot, minware can obtain granular agent session data via OpenTelemetry, which is supported at all plan levels (as of the publication date) with local file-based configuration. Additionally, certain plan levels of Claude Code, Codex, and GitHub Copilot allow central settings management to make rollout easier, though central administration is always possible with mobile device management (MDM) tools as well. Cloud agent sessions are also traced with OpenTelemetry, which can be configured via a vendor interface for vendor-managed agents, or via environment variables with self-hosted cloud agents.
For Cursor, an enterprise-level subscription is required (as of the publication date) for minware to access enterprise-only APIs and obtain agent session data for both cloud and local agents to calculate the session automation metrics listed in this pillar. Cursor users without enterprise can still see overall token spending and usage per person and day, as well as other metrics in this pillar outside of session automation.
The human intervention rate and automation runtime metrics further utilize an AI classifier. For enterprise minware customers, a forward-deployed engineer (FDE) can implement and maintain a classifier optimized for the organization’s environment. minware Pro customers can directly customize a default classifier by adjusting labels and examples in a self-service interface.
Ticket Work-in-Progress (WIP)
The efficiency goal of the operational outcome pillar left off looking at ticket cycle time, but ticket cycle time alone can be difficult to debug and identify opportunities for improvement.
The ticket work-in-progress (WIP) metric drills down one level to expose inefficient workflow steps at the inter-person level, as well as delays in the release process.
The nice thing about ticket WIP is that it gives you an absolute measure of performance because the optimal value is close to 1. A single person working uninterrupted with efficient automated tools will never be blocked and always have a single ticket in progress. Ticket WIP values much greater than 1 indicate that work is piling up on a person who is overloaded, a developer is waiting on a slow automated process (potentially a CI/CD job or an agent), or there are interruptions from bugs or urgent unplanned work.
minware measures ticket WIP by default by aggregating all the time intervals each day where a person is assigned an in-progress ticket and dividing that by a day. This metric can be customized to aggregate by team if people often work together on tickets, or adjusted to treat different statuses as in-progress.
In an analysis of anonymized customer data, we found people in the lowest quartile of ticket WIP came in below 1.5, making it a realistic target that we recommend. A WIP of 1.5 allows some overlap to start on the next task while waiting for code review or a deployment, yet still flags higher-WIP scenarios that cause significant context switching.
minware offers optional guardrails for ticket WIP that are not part of this framework, but can be useful for teams with inconsistent ticket hygiene. These guardrails ensure that ticket WIP does not show wrong values in the case of untracked or stalled work. The first is a “has in-progress ticket” metric that looks at whether each person has an in-progress ticket when they make each commit. This distinguishes between true idle time and time where someone is working on a task off-the-books. The other is “stale tickets,” which looks at tickets that have commit activity at some point, but go seven days without a commit or status change, indicating that they may have been improperly left open even though work has stopped.
Ticket WIP is the first canary in the coal mine for teams that are using AI to start on more work in parallel and losing productivity to context switching rather than investing in best practices that make each ticket faster.
Ticket Size
One way to make your ticket WIP 1 is to put everything into one big ticket for the whole sprint. While this is an extreme example (which we’ve actually seen, but it’s rare), any tickets that are larger than they should be will falsely show better WIP than the reality.
Measuring average ticket size directly and looking at outliers will help teams ensure that there is an effective up-front planning process. In addition to making workflow metrics more accurate, carefully thinking through the steps required to complete a task will drastically reduce the risk of missing requirements and experiencing rework or estimate misses down the line.
By default, minware defines ticket size as the ticket’s story point estimate. Alternatively, some customers configure ticket size to be the actual work time that went into a ticket according to minware’s time model, or they look at both. While story point estimates don't capture underestimation, they are available when work starts, giving teams the opportunity to re-think their plans.
To optimize this metric, we first recommend looking at any single ticket that represents a week or more of capacity for one person. For story point estimates, this is any ticket taking up half or more of a person’s two-week sprint capacity. Week-plus tickets should almost always be broken down into smaller tickets that are closed after reaching milestones even if a final deliverable isn’t ready until the end.
Next, we recommend an average target of two or fewer days of capacity for one person – one-fifth or less of a person’s two-week sprint capacity in story points. This indicates that tasks are consistently broken down into smaller pieces as a standard practice and independent changes are not being batched together. The two-day target is a general rule of thumb rather than a benchmark from customer data.
When splitting tickets, it is best to keep the total story points equivalent, using fractional points as necessary to avoid skewing other metrics like roadmap value delivery.
When working with AI, small tickets are particularly important because they leave less room for low-value runaway token use. As tempting as it may be to let agents run wild on large, open-ended tasks (“what if it just works!?”), our experience is that large tickets are a huge slop-generation trap that harms productivity and AI impact.
Pull Request (PR) Complexity
Breaking down work into small tickets at a functional level is helpful, but even small units of functionality can require several distinct code changes that should be broken down into separate pull requests.
In a recent article, the Pragmatic Engineer highlights “produce less code” as a way to combat code review overload, suggesting PRs less than 500 lines of code, splitting up PRs, and using stacked diffs (which makes splitting PRs easier).
Every change that is wrapped into a single PR when it could be separated (such as refactoring code before modifying it) makes the PR riskier and more difficult to review in a superlinear manner – the combined PR is higher risk/effort than the sum of the split PRs. This happens because the reviewer has to sift through more changes and think about more potential interactions between different components.
The pull request (PR) complexity metric reflects the risk and review effort associated with each pull request. minware measures PR complexity based on the number of added code lines by default, but the metric can be customized in various ways, such as excluding test files and fixtures or counting the number of files modified. Deleted lines are omitted by default because they usually contribute far less risk than new lines.
Truly measuring the cognitive burden of a pull request is notoriously difficult. A small change to a module with many dependencies may be harder to reason about than a huge change to a simple form used on a single page.
This is why we recommend using PR complexity as a first step to create a list of potentially problematic PRs, and then having a senior engineer investigate them further. We advise starting with the 500-line threshold suggested in the Pragmatic Engineer article against the default lines-added definition, and then adjusting it up or down depending on how many PRs above the threshold have a real issue. In this secondary investigation, the engineer should answer: could and should this pull request have been split up?
Assessing PR splitting is nuanced. One approach is to consider the current state of the code and determine whether the PR should have been broken apart. Pull requests that fail this test reflect an engineering mistake – often rooted in problems with agent instructions or technical planning.
A broader approach for assessing PR splitting is to consider whether it would be possible to split up the pull request with improvements to the code base. For example, better test harnesses may enable thoroughly validating and releasing API changes in a separate PR from the UI that exercises them, even if that’s not practical currently.
On the other side of the assessment, PRs should generally not be split if they cannot be reviewed independently. Requiring a reviewer to switch back and forth between PRs to understand the scope and risk of a change just masks true PR complexity.
It is also important to note that splitting PRs to reduce complexity will mechanically improve metrics with PR count as the denominator like bug rate and change failure rate. Those metrics can be configured to use a different denominator to mitigate this effect.
Both splitting approaches yield action items that will help teams work in easier-to-understand code batches, which is critically important for improving AI impact because it relieves pressure on code review and decreases the risk of bugs.
Pre-Merge Churn Rate
Even if PR complexity is low, the final code in a pull request does not reflect the path it took to get there. A small PR with a single commit is far different from a small PR with thousands of lines added and then subsequently removed.
The pre-merge churn rate metric captures the amount of code that was committed and subsequently deleted before the pull request merged.
minware calculates the pre-merge churn rate as the total lines added that don’t make it to the final merge across all commits divided by the total including both churned and merged lines.
If a lot of code makes it to a commit boundary and is subsequently removed, this indicates a change in high-level code structure. Smaller bug or style fixes that touch a few lines but leave the structure intact do not cause much churn.
Pre-merge churn is great at isolating oversights with up-front architectural planning, open-ended prompts that lack specific instructions, and failure to discuss alternatives and solicit feedback from agents at the start of a session.
When optimizing pre-merge churn rate, we recommend focusing on pull requests with more than 50% churn as a rule of thumb and keeping the average below 30%. Changes and refactoring often produce some churn, which can be positive if it represents refinement rather than completely discarding work. These targets are also high enough that they indicate significant rework even for very frequent commits, such as after every prompt.
Since some churn is expected, we recommend digging into overall commit frequency for developers or teams with less than 5% pre-merge churn over a longer time period. This ensures they are committing regularly and pushing all their commits. (minware loads original commits for squash and rebase merges across all version control providers, so those will show up.) This check also catches the scenario of lax code review or inadequate changes to address reviews.
By default, this metric does not count PRs that are closed and never merged, which prevents it from counting intentional spikes and prototype PRs. The metric can be customized to measure closed PRs as 100% churn if they are a significant source of churn, in which case it can be helpful to add labels for intentional spikes/prototypes to exclude them from consideration.
Having an agent write a bunch of code only to heavily refactor or discard it later is one of the biggest sources of slop that hurts AI impact. Pre-merge churn rate is the key metric for driving down this form of waste.
Post-Merge Churn Rate
Even when everything goes smoothly up to the point of release, you’re still not out of the woods. Production deployment is the real test for code, and it may uncover problems missed during code review.
We already have bug rate and change failure rate to detect visible defects post-release, but those metrics don’t catch issues with code quality that may lead to heavy refactoring or deletion.
The post-merge churn rate covers this potential source of slop: escaped defects in code quality. It measures how many lines are removed shortly after being merged to a main branch. minware computes post-merge churn for each PR merged into a main branch by dividing the number of code lines merged that were removed within 14 days by the total lines merged, with each merge’s churn recorded at the end of the 14-day window (so the metric lags two weeks). This captures changes that are reverted or refactored within a typical two-week development cycle. minware calculates the overall metric based on the total number of lines instead of the PR count, so larger PRs will be more heavily weighted.
The window size is configurable, and customers may look at multiple windows, such as code that churns within a day, indicating an urgent revert. It is also important to exclude large programmatically generated and data files like test fixtures from this metric. These exclusions can be shared with those for PR complexity, or configured separately. minware can also break down churn by repository, path, or file extension to show where churn is happening.
Because churn can vary wildly and only some of the reasons for churn are problems in need of remediation, we suggest drilling into data points aggregated by developer and week that have over 15% churn to obtain a short list of large-churn PRs for further investigation. This target is based on an assessment of our customer data and churn levels that are likely to reflect code quality problems rather than typical iteration. This metric is intentionally not grouped by team because aggregating over more people averages out variance and makes the 15% threshold less sensitive as team size grows. However, minware does show the number of people above the threshold to provide a team-level view.
This metric is weighted by line count, so we advise looking at PRs in descending order of absolute churn line count when drilling into a data point. For each high-churn PR, an engineer should assess whether the churn reflects a significant structural code quality problem that should have been pre-empted with better up-front technical planning and slipped past code review. If so, the engineer should determine the root cause of the oversight.
We consider post-merge churn a best practice rather than an operational metric because unlike bugs and change failures, not all churn is bad and making that judgment requires interpretation by the team doing the work.
Human Commit Rate
When looking at the efficacy of agent sessions, the first question is whether engineers decide to use them in the first place. People may still write code manually if they don’t believe an agent will make it easier or the agent is lacking tools or access to complete the task.
The human commit rate metric looks at the number of commits that are not traceable to an agent session. minware computes this metric by looking at the number of original (non-rebased) commits that fail to link to an agent session through the hypercube data model and dividing by the total number of original commits.
This metric additionally highlights gaps in data coverage by breaking out commits that show evidence of AI use in the commit message, but do not match AI usage data for that user.
We recommend customizing this metric to exclude sources of commits where agent use is not expected or intended. minware discounts bot and merge commits from version control systems by default, but other commits like small configuration changes or documentation updates may make sense to exclude as well.
When optimizing this metric, teams should set a target that makes sense based on the nature of their work and the maturity of their best practices. Modern agents have become very effective at writing code, so a steady-state rule-of-thumb target is under 10% overall, and under 1% for commits with evidence of AI use that fail to link. Unexpectedly high human commit rates should be investigated for potential data issues as a first step because missing telemetry will inflate the rate. The break-out of commits with AI evidence may not detect all missing telemetry because some agents do not leave markers in commit messages.
To identify improvements, teams can regularly sample a fixed number of human commits (e.g., 20) to assess whether they represent a problem with agent capabilities, or whether they were just short and easy enough for a human that AI use was not necessary.
Agent Session Length
Now that we’ve passed the gauntlet of metrics to catch problems with workflow, planning, and prompting, we’re left with agent sessions that should ideally be short. Yet, a variety of pitfalls remain that can sink AI productivity even with perfectly crafted prompts and granular PRs.
Agent session length is a powerful metric because it covers the myriad ways agents struggle on the path to task completion.
minware measures agent session length directly from start and end times found in AI tool session data.
Customers can configure agent session length to split up sessions that are open for a long time or cross pull request boundaries. However, those may still be of interest because they can surface anti-patterns with relying on session context rather than integrating it back into agent instruction files.
Certain integrations like OpenTelemetry with Claude Code provide further detail on how much time the agent was actively working in a session. It can be helpful to break down agent session length by agent vs. human activity to isolate issues with slow automated steps from issues that manifest in human intervention.
To optimize agent session length, we don’t recommend a default target because session length can vary wildly depending on each team’s maturity level with all of the best practices in this pillar. Instead, we suggest that each team look at a fixed number (e.g., 10 or 20) of the longest sessions each week to assess whether their length reflects a significant issue that requires remediation. When the longest sessions become mostly free from best-practice violations, the team can then establish steady-state targets that align with their work patterns to prevent regressions. Due to the long-tailed nature of sessions, we recommend tracking the median, 90th percentile, and a high-90s percentile (the exact one that makes sense depends on the number of sessions in each time period) rather than using an average, and setting targets for those numbers once they stabilize.
Sessions that are short may also reflect an agent failure and human takeover. However, it is not necessary to look at short sessions with this metric because that type of problem is covered by the human commit rate metric.
While agent session length alone won’t tell you why sessions are long, it is a core gauge of AI impact best practices covered in the sections that follow.
Human Intervention Rate
An agent session contains a series of instructional prompts telling the agent what to accomplish. In an ideal scenario, each prompt would start with “Thanks Claude, that was perfect, here’s the next step…” Or, the entire session would be one shot from a well-structured plan. In reality, this ideal is so rare that most AI users would consider it cause for celebration.
In lean methodology, hand-offs between people are a major productivity killer. With AI, hand-offs between agents and people degrade productivity in the same way: the other party has to stop and wait every time the session changes hands. For the person, this can add significant context switching overhead.
The human intervention rate looks at the count of permission escalations and other prompts in between instructional prompts divided by the number of instructional prompts in the session.
minware calculates human intervention rate with a custom-trained AI prompt classifier, which can be tuned self-service or maintained by a forward-deployed engineer for enterprise customers as described above.
Of all the metrics in the framework, human intervention rate does the most heavy lifting for day-to-day identification of areas that need improvement. It can uncover a variety of issues, but these are the most common ones detected by this metric and their corresponding remediation actions:
| Intervention Type | Description | Remediation |
|---|---|---|
| Permissions Escalation | The agent asked for approval to execute a tool or command. | Consider creating a locked-down container to safely allow disabling permission checks, or build out allow-lists of known safe commands and add agent instructions to avoid commands that trigger escalations. |
| Agent Convention Violation | The agent did something that was incorrect according to the conventions of the repository or failed to execute a required step like running tests. | If guidance is missing, add it to an agent instruction file (e.g., AGENTS.md, CLAUDE.md). If guidance is present but not followed, improve it or consider splitting the main instructions file into smaller task-specific files to prevent context bloat. |
| Human Manual Testing | The human performs testing outside of the session and provides feedback on output to the agent. | Add missing test harnesses to enable the agent to write automated tests that do full end-to-end validation. Add tools and MCPs to give the agent access to all the systems needed for testing. Improve sample data sets and sandbox environments to enable self-contained testing without accessing production data/systems. |
| Human-Directed Testing | The human directs the agent to perform some testing with one-off scripts or a tool that is not part of the automated test suite. | Improve prompts and agent instructions to have the agent write necessary automated tests as part of the task instead of doing ad hoc testing. |
| Human Information Lookup | The human retrieves some information that the agent cannot or does not know how to access. | Invest in MCPs, tools, sandboxing, anonymized data sets, and security controls that let the agent safely access the information it needs. |
| Agent Off-Track | The human sees the agent going down a bad implementation path and has it change direction. | Improve up-front technical planning and provide better guidance in instructional prompts. |
| Agent Mistakes | The human gives the agent feedback on mistakes it made implementing instructions. | Convert non-deterministic LLM tasks to deterministic code where possible. Add comprehensive automated static checking tools if missing. Improve prompts and agent instructions to avoid common mistakes. Consider adding code review tools or external review agents. |
Permissions-escalation interventions can be trivially reduced by dangerously loosening security controls, so be clear in your guidance that only safe remediations should be made when optimizing this metric, like implementing a sandbox or avoiding dangerous commands entirely.
Much like agent session length, the human intervention rate is going to vary wildly based on the maturity of best practices and difficulty of the work in each session. As such, we don’t recommend a default target. Instead, we suggest looking at the most prevalent intervention type labels and investigating sessions with the highest rate of those interventions. This approach isolates root causes that will have the biggest impact when fixed.
Once each team reaches a steady state where issues found in sampling are no longer worth fixing, then the team can establish target intervention rates. It can be helpful to segment these by intervention type if certain interventions are a lot more common than others.
Another failure mode with intervention rate is for people to just not intervene when they should. We don’t recommend addressing that issue with this metric because it should be caught by higher-level metrics like bug rate, or manifest in high pre-merge churn if the problems are detected and addressed during code review.
As you can see from the length of this list, remediating human intervention is where individual engineers and teams will spend a large part of their time optimizing AI’s impact on end-to-end value delivery.
Cloud Agent Use Rate
Once a team has established good planning practices that shorten agent sessions and enabled their agents to work with minimal human intervention, the final frontier is running agents entirely in the cloud. This typically involves kicking off an agent session from a tool like Slack, Linear, or Jira, and culminates in an open pull request.
The cloud agent use rate measures the portion of agent sessions executed in the cloud. It reflects the team’s confidence that an agent can successfully complete or make progress on a task entirely on its own.
minware computes the cloud agent use rate by counting the number of sessions started with a cloud agent, divided by the total number of cloud and local sessions.
As cloud agent use rate starts to increase, it is helpful to look at remaining local sessions and assess whether something is holding them back from the cloud. It may be technical limitations of the cloud environment, or simply an expectation of required intervention.
Some remaining local sessions may not be suited for the cloud, such as interactive planning. The goal for cloud agent use rate is not 100%. We recommend sampling remaining local sessions as cloud agent use rate goes up and setting an appropriate target like 80% based on the nature of the team’s work.
The cloud agent use rate is a helpful leading indicator of AI maturity because it shows that the team is confident in AI completing the entire development process without human assistance.
Cloud Agent Success Rate
While deciding to use a cloud agent gauges the team’s confidence, it’s important to pair that with a measure of whether those sessions actually succeed.
A successful cloud agent run means that the agent can execute the entire code workflow without any setup or cleanup work from a human developer.
minware computes the cloud agent success rate by counting the number of cloud sessions that are free from a reported error, do not have human or non-cloud AI commits at the end of the pull request, and whose pull request is merged. Then, it divides that by the total number of cloud sessions.
This metric can be customized to extract the success or error status from agent telemetry or other indicators on the pull request or ticket. It can also be configured to treat sessions where the cloud agent made changes following a human PR review as failures. minware also breaks down failures by reason or error code.
Cloud agent success rate validates that AI consistently completes work without failures or human intervention. When paired with cloud agent use rate, it is an excellent barometer of AI impact on value delivery because it shows that the entire development process has been enabled for AI to work without human assistance.
Note on Cloud Agents and Work-in-Progress
Cloud agents can increase the total amount of work in progress at the team level. However, because they execute at the beginning of the task before an engineer has picked it up, they do not drive up human WIP and therefore should not be counted as part of the ticket WIP metric, which is designed to address human context switching. Cloud agents should ideally reduce human WIP because they shorten the time that a person has to work on a task, which begins later at code review.
While parallel cloud agents that queue up open pull requests don’t increase human WIP, any cloud agent work that happens after receiving a code review to address comments is in the human workflow and liable to cause context switching, so should be counted toward ticket WIP. We recommend excluding cloud agent work from ticket WIP with ticket statuses that indicate the agent is working on a task prior to engineer involvement.
Also, a long queue of open pull requests can still be harmful even if it doesn’t increase human WIP because it can both overwhelm the human review bottleneck and lead to merge conflicts. While being able to run cloud agents is beneficial and captured by the cloud agent use rate metric, it’s important to keep in mind that cloud agents should only be kicked off when there is available capacity to review their work.
Finally, it is still better if cloud agent sessions are as short as possible for the overall workflow and ticket cycle time, and this is captured by the agent session length metric above.
Automation Runtime
At this point, we’ve gone through a battery of metrics to ensure that agent sessions autonomously produce correct output. There’s one final issue that can majorly affect AI impact: slow automated tasks.
The automation runtime metric accounts for all the time that agents spend waiting on skills, tools, commands, and MCPs, including CI/CD builds and automated tests.
minware collects tool use information from granular AI tool session data. minware computes the total and average runtime along with total invocation count grouped by skill, tool, and MCP name.
This metric can also be configured to categorize unstructured shell commands using an AI classifier, which can be tuned self-service or maintained by an FDE for enterprise customers as described above.
As a rule of thumb, we recommend targeting under 30% of active agent session time as a budget for total automation runtime. Pushing much below that has diminishing returns as model API requests dominate overall latency. minware computes active agent session time by pulling it directly from OpenTelemetry reports for Claude Code, and approximating it using a heuristic based on prompt timing for other tools.
Non-AI tool runtime is often overlooked by teams focusing on AI impact, but it can represent a major component of overall task length. It is especially important because agents tend to execute tools like automated test suites far more often than developers did before AI.
AI Rules of Engagement
The final best practice in minware’s Lean AI Framework is the most crucial for overall success. Even if an engineering leader rolls out every metric here, value can fail to materialize if the humans who participate in the software development process don’t buy in.
To truly stop slop and create value, everyone in the organization must take responsibility for the quality of their AI-assisted work output, and be held accountable if they shirk this responsibility.
The AI rules of engagement best practice involves establishing a written artifact that everyone on the team has read, had the opportunity to critique, and agrees to follow when working with AI.
We recommend measuring this best practice outside of minware with a binary check: does an AI rules of engagement document exist and has it been reviewed by each team in the last year?
Depending on the culture of the organization and the mix of seniority levels on each team, expectations may vary about the level of quality and self-review each member should uphold. However, the goal is universal: each person should feel like their peers respect their time and dedicate sufficient effort to not send them slop.
AI rules of engagement are successful when each person perceives that they spend more time on stimulating, value-creating work with the addition of AI than they did before.
The best way to assess AI rules of engagement and the cultural impact of AI is for managers to discuss how each of their direct reports feels about the topic in one-on-one conversations, and for leaders to ask about it in skip-level discussions.
Surveys can also be helpful (minware doesn’t offer surveys but can integrate with third-party survey tools), but personal conversations are essential for getting at the truth of sensitive issues like whether coworkers respect you, or, in skip-levels, whether people feel supported by their manager.
Avoiding Metric Gaming
Introducing a measurement framework can raise concerns about gaming metrics.
Before addressing those concerns, it’s important to recognize that bias may exist without metrics. The goal of metrics should be to reduce bias by giving managers quantitative data points to confirm or refute their initial opinions.
The first and most important principle we advocate is for managers to always hold the final judgment about performance. Metrics should augment that judgment rather than supersede it. Managers may elect to use certain metrics as part of performance evaluation, but that should be at their discretion. They should also make it explicit to their direct reports before any metrics are used for evaluation: (1) which metrics are being used, (2) that no additional metrics will be used without disclosing them first, (3) how their values will be interpreted, and (4) that the manager reserves the right to not use any metric, falling back on other assessment means if the metric is no longer in alignment with its intent. These rules remove much of the incentive to game metrics without removing the incentive to optimize them and establish trust between individuals and their managers.
In the context of the pillars in this framework, the highest reporting level listed by each pillar should also serve as a guide for avoiding abuse. The reporting level does not mean that higher leaders should never see the metrics below. Instead, those metrics require interpretation, context, and presentation from their owner. Leaders should not form any conclusions based on lower-level metrics absent this presentation.
Teams in particular should never have to worry about someone looking over their shoulder from the outside and judging their best practice metrics. They may share them where relevant to highlight progress or show what they’re doing to improve operational metrics, but the decision about whether and how to share should be theirs.
For the operational and business outcome metrics that a leader may judge without the team in the room, it is critical to have clear definitions and auditing. minware makes this easier because the exact metric formulas and customizations are visible in the application.
We recommend having engineering managers review operational and business outcome metrics tied to their team on a quarterly basis. As part of this review, they should suggest any changes to metric definitions for their team, list any outliers that deserve exclusion, and provide footnote explanations for things that are likely to raise questions. This gives teams the opportunity to make sure they are judged fairly, and gives leaders visibility into different metric definitions and outlier exclusions on each team.
Finally, if leaders have lingering concerns about metric gaming, they should resist the urge to casually browse team dashboards and instead conduct a structured audit similar to a security or financial audit. The audit should outline a process for collecting specific evidence for a random sample of teams, people, projects, tickets, bugs, sprints, and pull requests. The audit may verify that operational and business outcome metrics are free from inaccuracies caused by data problems, improper metric definitions, or practices that manipulate the metric’s accuracy, whether intentional or not. The evidence collected and initial findings should be shared with each manager, and the manager should have the opportunity to respond to any adverse findings. A transparent and methodical audit process is essential for maintaining mutual trust and accountability between teams and leadership.
Conclusion
This article introduces minware’s Lean AI Framework, which provides a comprehensive roadmap for engineering leaders, managers, and individual contributors to accelerate value delivery by cutting down on slop.
This novel approach of focusing on slop reduction is rooted in long-established principles of lean methodology. It is an antidote to the output-first mentality of previous measurement frameworks.
The framework’s metrics depart from the off-the-shelf model of previous frameworks. They support extensive customization to uniquely align with how each organization works, making them well-suited to distinguish between slop and value in a way that generic metrics cannot.
minware’s framework also includes an extensive set of best-practice metrics that pick up where other frameworks leave off: actionable guidance about how to improve AI impact.
Finally, the framework looks at the human aspects of AI impact and offers guidance for managers and leaders about maintaining alignment and trust during the process of AI adoption.
The most important thing to remember when rolling out this framework is that it’s not just about AI impact, it’s about using AI to maximize human impact.
To get started, kick off a free trial: