top of page

Is AI Making Organizations More Effective, or Just Busier?

5 hours ago
7 min read

We may be repeating some familiar mistakes in how we deploy AI, measure productivity, and define success.


There is something strangely familiar about the current rush to deploy artificial intelligence across organizations.


We are counting licenses, tracking adoption, measuring how frequently employees use AI tools, and celebrating increases in individual productivity. Executives are being presented with dashboards showing how much code was generated, how many documents were summarized, and how many hours employees supposedly saved.


It all sounds remarkably like the early days of enterprise Agile transformation.


Back then, we counted Scrum teams, certified practitioners, velocity, and the number of people trained in new ways of working. These measurements gave us the appearance of progress. Yet they frequently told us very little about whether customers were receiving better outcomes, whether organizations had become more adaptable, or whether we had actually improved how value flowed through the system.


I worry that we are making the same mistake with AI, only at a much greater speed.


We are measuring how much AI is being used rather than whether it is making our organizations more effective.


And that distinction matters more than many organizations realize.


The Productivity Illusion


Let's consider a relatively simple example.


A software development organization introduces AI coding assistants. Developers quickly discover they can generate significantly more code in less time.


Management celebrates. Productivity is up! The investment is working!


But what happens to the rest of the organization?


Testing teams must now validate substantially more code. Security reviews increase. DevOps teams face additional integration and deployment demands. Support teams must absorb the consequences of whatever reaches production.


The developers have become faster, but the organization hasn't necessarily improved its ability to deliver value.


In fact, it may have made things worse.


Imagine that a development team previously generated 40 units of work per week, roughly matching the capacity of its downstream delivery system. After adopting AI, developers can produce 100 units, while testing, integration, and deployment capacity remains at 40.


We haven't increased organizational throughput. We've created 60 additional units of unfinished work every week.


This is a hypothetical example, not research data. It assumes comparable work items, unchanged downstream capacity, and no other constraints.


More work in progress. Longer queues. Greater coordination overhead. Potentially more defects and frustration.


Yet the development productivity dashboard looks fantastic.


This is a classic systems thinking problem. Optimizing one part of a system does not guarantee an improvement in the whole.


AI hasn't eliminated the constraint. It has simply moved the pressure somewhere else.


Are We Measuring the Wrong Things?


Much of today's AI measurement appears focused on what is relatively easy to count.


How many employees are using AI? How frequently? How many prompts are submitted? How many hours have supposedly been saved? How much content or code has been generated?


These measurements aren't inherently useless. They can tell us whether technology is being adopted and where activity is increasing.


But adoption is not effectiveness, and activity is not value.


We should be asking a different set of questions:


Instead of asking...

We should also be asking...

How many people use AI?

What meaningful outcomes have improved?

How many hours did AI save?

Did overall delivery time decrease?

How much more code was generated?

Did reliable software delivery improve?

How many tasks were automated?

Did we eliminate unnecessary work?

How much did labor costs decrease?

What happened to total cost and organizational capability?

How many AI agents were deployed?

Did coordination become simpler?


One particularly interesting warning came from a 2025 METR study involving 16 experienced open source developers completing 246 real tasks.


Participants expected AI tools to reduce their completion time by approximately 24%. Even after completing their work, they believed AI had reduced their time by about 20%.


However, the researchers measured something quite different. In that study, developers using AI took approximately 19% longer to complete their assigned tasks.


The findings were specific to the tools, participants, and tasks studied. In a February 2026 follow up, METR said newer tools likely offered greater speed improvements, but selection effects and unreliable time reporting prevented a dependable estimate. The earlier result should not be generalized to all developers or current AI systems.


Nevertheless, the study illustrates an important problem: perceived productivity and actual productivity aren't necessarily the same thing.


And even actual individual productivity is not the same as organizational effectiveness.


The Hidden Costs We Aren't Counting


Another concern is that AI may not eliminate as much work as we think. Instead, it may change the nature of the work and where the effort occurs.


Someone still needs to verify the output, understand the context, resolve inconsistencies, integrate the results, and accept accountability for the outcome.


Consider five areas that deserve much more attention in AI effectiveness measurement.


1. Verification burden


How much human effort is required to review, correct, and validate AI output? Are we eliminating work or simply converting creation activities into verification activities?


2. Coordination cost


If everyone can generate substantially more information, decisions, documents, and code, who integrates it all? Are we reducing organizational complexity or increasing the amount of work required to maintain shared understanding?


3. Cognitive debt


What happens when people increasingly accept AI generated recommendations without understanding the underlying reasoning? Are we improving organizational intelligence, or gradually weakening our ability to solve problems independently?


4. Capability development


Many entry level activities serve an important purpose beyond their immediate output. They help people develop experience, judgment, and expertise. If we automate those activities away, how will we develop the next generation of senior engineers, analysts, designers, and organizational leaders?


5. Organizational adaptability


Perhaps most importantly, is AI improving the organization's ability to recognize changing conditions, make informed decisions, and respond effectively?


These are proposed assessment dimensions, not standardized AI performance measures. Each needs concrete measures and validation in the setting where it is used.


These are not simply technology questions. They are questions about how organizations function and evolve.


Are We Automating Dysfunction?


There is another trend that deserves scrutiny.


Organizations are introducing AI into existing operating models without sufficiently questioning whether those models make sense in the first place.


We automate reporting rather than questioning why the reports are necessary. We use AI to prepare for meetings that perhaps shouldn't exist. We accelerate approval processes without examining why so many approvals are required.


We build intelligent agents to navigate organizational complexity instead of reducing that complexity.


The result?


We risk making dysfunctional systems operate faster.


This is where I believe the Agile and Lean communities have something particularly valuable to contribute.


For years, we have talked about understanding value streams, reducing unnecessary work, making constraints visible, improving feedback, and creating organizations capable of learning and adapting.


Those principles haven't become less relevant because of AI.


Quite the opposite.


The technology has made them more important.


Before asking how AI can accelerate a process, perhaps we should first ask whether the process needs to exist.


Before automating a task, we should understand the value it contributes.


And before celebrating productivity improvements, we should determine whether they have improved the effectiveness of the entire system.


A Different Way to Measure AI Effectiveness


I would suggest evaluating AI deployment across four connected dimensions.


Task effectiveness: Did AI make a particular activity faster, less costly, or more accurate?


Flow effectiveness: Did those improvements translate into better end to end delivery? Are we reducing lead times, unnecessary work, rework, and constraints?


Business effectiveness: Did the organization realize meaningful improvements in customer outcomes, financial performance, quality, or its ability to respond to emerging opportunities?


Organizational effectiveness: Are we developing a more capable, adaptable, resilient organization, or are we trading long term capability for short term efficiency?


The distinction is important.


An organization can achieve tremendous improvements in task effectiveness without experiencing meaningful improvements in business effectiveness.


Likewise, short term financial savings may look attractive while gradually undermining the knowledge and human judgment the organization will need in the future.


Success at one level doesn't guarantee success at the next.


Microsoft Research's New Future of Work Report 2025 also points toward collective productivity: improving how teams, organizations, and communities work together, beyond individual gains.


Start With Value, Not Technology


Too many AI deployment strategies appear to follow a predictable sequence:


Buy the technology. Drive adoption. Automate tasks. Measure usage. Demonstrate productivity. Look for a return on investment.


What if we reversed that thinking?


Start by identifying the outcomes we want to improve. Understand the existing system and its constraints. Establish meaningful baseline measurements. Then experiment with AI where it has the potential to make a difference.


Observe what actually changes, including unintended consequences.


Learn. Adapt. Scale what works.


And sometimes, the most valuable discovery might be that AI isn't the appropriate solution.


Perhaps the process should be eliminated. Perhaps a policy needs changing. Perhaps the organization needs fewer handoffs, clearer accountability, or better collaboration.


Technology is only one possible intervention.


What Must Remain Human?


Finally, I think we need to be careful about how we define efficiency.


Not every activity that takes human time is waste.


Learning takes time. Developing judgment takes experience. Building relationships requires interaction. Challenging assumptions, navigating uncertainty, and developing shared understanding are fundamental to how organizations learn.


We should certainly explore how AI can support and enhance these activities. But eliminating human involvement simply because automation is technically possible may create consequences we don't fully understand until much later.


The goal shouldn't be to remove humans from as much work as possible.


It should be to create systems where humans and AI contribute in ways that improve our collective ability to deliver value, solve problems, and adapt.


That requires considerably more thought than counting licenses or estimating hours saved.


The Question We Should Be Asking


I am optimistic about the potential of AI. The opportunities to reduce unnecessary effort, expand our capabilities, and rethink how work gets done are substantial.


But I am increasingly concerned about how we are defining success.


We've been here before with other transformation movements. We focused on adopting practices, implementing frameworks, and measuring activity, sometimes losing sight of the outcomes those changes were intended to produce.


Let's not repeat that mistake.


AI adoption is not AI transformation. Individual productivity is not organizational effectiveness. And generating more output does not automatically mean creating more value.


Perhaps the most important question leaders should be asking today isn't:


"How much more can we produce with AI?"


It's:


"Are we becoming a more effective organization because of it?"


Those are two very different questions. And the answers may surprise us.


What are you seeing in your organization? Are AI investments improving the flow of value, or are they creating new bottlenecks and hidden work? I'd welcome your thoughts.



Research references




 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
About nuAgility
nuAgility is a consulting and community-driven organization focused on helping companies and practitioners improve how work actually gets done.  Through hands-on engagement and open community conversations, we explore and teach practical ways to deliver value in complex environments.
Take the Next Step

Organizations

Improve how work actually gets done across teams and systems.

​

See how we help reduce complexity, align work to outcomes, and build more adaptive organizations.
 

Practitioners

Grow your ability to navigate and shape real-world work.

​

Explore insights, tools, and learning experiences designed to move beyond theory into practical application.
 

Community

Engage with others working through the same challenges.

​

Join open conversations with practitioners sharing real experiences, ideas, and lessons learned.
 

bottom of page