Agentic AI: Hype or the Way to Unlock Productivity Gains in Software Engineering? – Predict 2025
In the last couple of years, AI assistants have gained widespread attention and many companies have rolled them out across their teams. However, reports on productivity impacts have been mixed. 2025 could be the year of agentic AI, as assistants will evolve into agents able to reliably perform larger tasks on their own.
In this session, you will learn about the developments in agentic AI, their differences to AI assistants, and what could make them a game changer in reaping productivity benefits.
We will focus on applications in software engineering where AI agents have the potential to increase coding efficiency, achieve continuous testing, improve code quality, and automate routine tasks.
This session is ideal for technology leaders looking to leverage AI for enhanced productivity and innovation in their software development workflows.
You will leave with the following takeaways:
– Streamlined interactions between the human and the AI agent are crucial for their benefit and success.
– Some use cases are already ripe for agentic AI today, others are less mature.
– Ultimately, true productivity gains come from autonomy.
Transcript
Hi, I'm Peter Smore and I will talk about agentic AI hype, or the way to unlock productivity gains in software engineering. Today you will learn about the path from AI assistance to agents. We look into use cases, maturity of agents and processes for using them to gain productivity in software engineering.
Finally, we will dive into an example of how to use agents for unit test generation. Some words about myself. I started my career as a software architect in the field of, uh, internet of Things IOT.
Then I did a PhD in static analysis, verification of embedded systems. I then continued my research at, um, university of Oxford, became a lecture in computer science, and then, uh, co-founded the AI for codes of Startup. And as a fun fact, I'm a keen back current here.
So in the last couple of years, AI assistance gained widespread attention and many companies have rolled them out across the teams. However, reports on productivity impacts have been mixed so far. 2025 could be the year of agent ai as assistance will evolve into agents able to reliably perform large tasks on their own.
Let's look at the common definition. An AI agent is an autonomous intelligent system that performs tasks without human intervention. Well, we don't really know what intelligent exactly means.
So the core of this definition is autonomous and human without human intervention. So we'll see IT. Assistance and agents support different use cases.
So as systems are useful in the course of the creative process of developing something, the usage is highly inactive and iterative in drafting solutions to small tasks. Also, they don't really need to produce perfect results all the time because the user can immediately fix the output at at low cost. So, however, in uh, autonomous agents, they require little to no induction.
There's the scale to large tasks, which of course require them to be most more trustworthy for the outputs to be correct. So we'll see that there is actually a continuum from assistance, uh, to agents. So the required level of precision depends on the use case.
So sometimes good enough is 80%, sometimes it means a hundred, a hundred percent. So if you want to have a fully autonomous AI agent, then it needs to be trusted a hundred percent of the time, whereas an AI assistant can only really deliver value when they are only right 80% of the time or even less so. The car industry has experience in developing agent systems for a long time.
The Society of Automotive Engineers, SAE, has developed a maturity model for driving automation, which consists of six levels in levels zero to two. The user, which is the car driver, is in charge of driving the car and is assisted by various automation features in levels three to five. The user is not driving the car, but the car is actually driven by the AI system autonomously.
Still at level three, the system might ask the driver to jump in when it can't handle the situation. So this is my take on a maturity model for AI systems. We have a couple of levels, um, and we distinguish these levels based on properties such as TEX initiative, which kind of fast to handle unit action, whether the system has actually the authority to execute for real, um, that require accuracy and ability to adapt and, uh, improve over time.
So level zero would be manual work. Then on level one we have assistance on level two. We delegate work to an agent on these two levels.
The initiative comes from the user, which is usually the case for doing some creative work. So assistance asset can handle smaller tasks, whereas to an agent that can delegate also larger pieces of work, the interaction with an assistant is in a tight, iterative interactive loop. And since it's supervised by the user at all times, the required accuracy for the assistant is not too high.
Of course, I'm going to be more productive with a more reliable agent. The larger the piece of work delivered by an agent, the more expensive it is to understand and thoroughly check its output, amend and rework it manually. So generally the more autonomous the agent, the higher the power for accuracy.
Hence, already a delegative agent is required to be much more reliable than new system. A delegated agent might also ask clear questions about the task to the user, but such interactions need to be minimized because otherwise the productivity benefit is lost by constant context. Switching on the user side.
On level three and form, we have agents that listen to the environment to take actions on their own, whereas the level three proactive agent may still ask questions to the user and require approval. The level four autonomous agent doesn't require any interaction and can act without you on approval. Of course, to trust an agent with such permissions requires perfect accuracy.
Maintenance tasks are typical examples for proactive agents where the agent identifies the needs to perform the task and delivers the results to the user for approval. The accuracy bar for such unsolicited work is higher than in the delegation case because the user needs their be weekly convinced about the utility and the quality of the work, and has little motivation to rework the output. The right most column is traditional automation like we used, for example, in continuous integration and delivery systems of CI/CD.
So traditional, uh, traditional automation is very similar to autonomous agents, but are dis is is distinguished by the nature of the tasks that it can handle, which is usually very well defined and and specific. So the, uh, also the traditional automation is usually has no ability to automatically adapt and improve all the time. So the cost for adaption is expected to be much higher for such a, uh, regional, uh, automation system.
So why and, uh, autonom reliability so important for unleashing productivity gains. So in this chart, we compare the difference in index actions when using an AI system to handle a certain task. On the horizontal task, uh, axis, we have the time and the different, uh, colors, uh, distinguish different kinds of, uh, parts of the task that need to be done to achieve, uh, the task.
For example, when I use an assistant, I need to first instruct the tool, uh, what it needs to do. Then I wait a bit for to receive an output, and if the result is not acceptable, then I have to rework and fix it and finally check and approve the work. So when I deliver a piece of work manually, then most all of the time is attended, which means a per person is actually sitting there and working towards delivering the task.
So with an assistant, they work can potentially be achieved, uh, faster, but the way the work achieved is achieved is much, uh, is very different. So it usually, uh, consists of multiple cycles of interactions with, uh, the tool. And, um, all this time is again, is is fully attended.
So the advantage of using an agent is that a big portion of the time is actually unattended, which means I can go away, um, do some other work, uh, in the meanwhile, but if the agent is unreliable and the work is is not acceptable, then I actually have to spend time in actually maybe reinst instructing the agent or even fully understanding the output and reworking it manually, which quickly results in the effort exceeding any productivity gains. However, with a reliable agent, it, the attended time will be tiny and the productivity gains can be traumatic. So with an autonomous agents, the at attend time will even be zero.
So agents only yield or or proportional productivity gains when handling rather large tasks and only unattended, asynchronous, uh, operation allows me to productively use the time while the tool is operating. So therefore interactions must be minimal to avoid determin context switching, uh, to the user. So assistance speed up part of the work, but, and, but everything is fully attended.
So the productivity gains through agents depend on their reliability to produce accurate results, or in other words, how much rework is acceptable so that there is still a productivity increase. The larger the output produced, meaning the larger the task handled by the agent, the more expensive it is to fix incorrect outputs. So only a reliable agents yield actual consistent productivity gains.
However, assistance are already useful with lower accuracy because I can fix output at low cost immediately, but the gains will also be not as big as for agents. So from what we have learned so far, there are two different effects on the productivity of software engineering that we need to understand. The first one is the so-called booster.
So let's take the example, uh, depicted in the upper half of this slide. So we need to fit, uh, three tasks. Task 1, 2, 3, enter a a sprint with, um, bounded capacity, let's say two weeks.
So here, task three doesn't fit in on the sprint, so we, we can't deliver it. However, if we have an, uh, automation booster, it'll shorten the delivery times for each of the tasks and we can actually fit it. So the booster helps us to deliver tasks by increasing, uh, velocity.
So another scenario is shown in the lower half. So here again, we need to fit three tasks. Task one and two have high priority, and task three has lower priority and is also quite large.
So typically such task will, um, never be deliberate because it, the higher priority tasks will always take precedence. So to deliver such a task, we we need an enabler to, uh, and which yeah, uh, which allows us to shrink task three to a reasonable size and fit it into the sprint and, and get it done. So difficult examples for such task like task three are, um, take that or cleaning up or modernizing the code base.
Such tasks will always take a backseat in comparison to features that deliver immediate value to customers. So having an enabler to tackle such tasks not only allow us to, to, to get them done, but also gives a secondary boost to delivery of, uh, actual customer features. So next, let's have a look at where there are automation use cases in the software development lifecycle.
This table shows various stages of the lifecycle from environments analysis to maintenance and and modernization. So environments analysis and validation are the hardest to fully automate because they require talking to the users together with design and planning. They are at the core of the creative process requiring a lot of human intervention interaction and integration.
So however, we can already see tools emerging there that for example, produce mockups of user interfaces based on our requirements. Then build integration delivery and deployment are already fully automated because they were the target of traditional automation endeavors through CI/CD pipelines. Currently the hottest area that receives most attention is implementation.
So we already have level one assistance there, like GitHub coot for example, and we are seeing the emergence of level two delegated agents. These agents are not as mature and can only handle small well-defined tasks. Uh, at the moment, use cases in other areas are more mature.
For example, in the area of of unit testing, there are predictive agents such as D Blue cover that automatically write tests on polygraphs and where they see it. Also for UI testing, there are agents such as fix AI that can automatically walk through user interface and find bug automatically similar. There are tools that can, uh, automatically pinpoint root causes of system fails in production, and there are tools that can perform complex maintenance tasks as automatically upgrading software depend is keeping the code base brain removing cold smells and fixing security vulnerabilities.
Today, none of these tools are really deployed in a fully autonomous way in the sense that they still require human approval to merge code into production. So gaining the necessary trust into this system takes some time, even if they're already operating at a very high, uh, at very high accuracy levels. So for example, the use case of unit generation for legacy systems can be considered level four autonomous, but enterprise company processes still require code path through review processes and approvals.
Uh, today in the last part of this talk, let's talk, let's look at, uh, agents for writing units. So I've identified here three use cases where we need to write unit tests. In the first one we have a untested code base, typically legacy base code base where we, the task is to write a regression unit test suite so that we can then make changes with, uh, more confidence.
The problem is that this is usually a multi-person new project and doesn't have high priority because it doesn't really solve any immediate, uh, customer problem. The solution is here, uh, to have an inhaler that take this problem like a level two or three or four agents such as deep blue cover that can write the regard, uh, tests fully automatically at scale, shrinking and on some mountable task inter effectively and brainer. The second use case here is writing new code for a feature where I need tests to check whether my implementation works fine.
So writing this test is not the most insurable task for developers and slows down the delivery of the feature here. A level one assistant can give us a boost. GitHub cover lot can do this, for example, but also the uh, digital covert inte login helps me.
The third case is continuous testing, where the goal is to maintain the level of testing of a code base. This involves fixing tests that start breaking, making sure the tests are written for all the code changes that make the way into the code base. The problems we are facing here are the lack of team discipline and also slowdowns in, uh, feature delivery.
Here, a level three agent such as deep blue cover can help to maintain the test automatically, uh, within the continuous integration system. So we conducted a study to compare coding assistance, uh, with agents. So you can think of that like a race between a assistant enhanced experienced developer and an autonomous agent to write unit tests for 20,000 lines of code within an available time box of 260 minutes.
At the end of this time, the autonomous agent in this case D Blue cover was able to achieve a coverage of 53%, whereas the GitHub colo enhanced, uh, developer only achieved uh, 15%. The main difference though is that they, uh, developer had to work hard for the whole time period, whereas the agent ran completely unattended and the developer could spend the time working on something else or just go for lunch on the business logic code for which uh, tests were written. The test quality measured by coverage and test strength was about the same.
The difference though is that, uh, the, um, agent was, uh, completely unattended. What makes this huge productivity differences happen is the higher reliability and accuracy 100% for the agent versus 64% for the assistant, which then in turn reduces the time in spent in injecting with the tool and reworking the results that the tool produced. So in the case of the agent here, no human induction is required apart from a few seconds for kicking off the tool.
This also allows for higher scale scalability. 4 million lines of code in a whole year, provided that they don't quit before that because of the sold This drawing drudgery that is involved in this task and agent in contrast can write tests for 25 million lines per year on a single laptop. And it is also cheap to scale up even further for parallelization because it's just compute.
So whereas with a coding assistant, one would require to use more expensive engineering stuff to scale. To summarize takeaways of this, uh, talk, so agents and assistants are already part of the engineer landscape today. Reliability and accuracy are crucial for smooth interactions and true productivity gains come ultimately from such reliable autonomous agents.
And unit testing is the most advanced use case on the agent material scale. And if you're interested in the deep blue carbon tool that I mentioned, we have recently released a developer edition for which you get a 50% uh, discount with this discount code. So thanks for your intention.



