Projects

“There still needs to be someone who can hit the brakes and say: This doesn’t make sense.” Petr Škoda on how GTS has evolved over the past six months

“There still needs to be someone who can hit the brakes and say: This doesn’t make sense.” Petr Škoda on how GTS has evolved over the past six months
Tereza Kopecká
Marketing Manager
15min

AI in testing is a topic that gets talked about a lot, but measured far less often. At Green:Code, we decided a few years ago to do things the other way around. And that is how GTS, the Green:Code Testing System, came to life. We spoke with Petr Škoda, who leads QA at Green:Code, about why GTS was created, how it has evolved over the past six months, and where AI still has its limits in testing.

What is GTS, and why did Green:Code decide to develop its own AI testing tool?

To put it in one sentence, GTS is a platform that uses AI to make end to end testing activities faster and more effective.

Why did we start developing it? The motivation was simple. For a long time, we struggled with limited QA capacity and overloaded QA teams across our projects. Our approach was often reactive, and we simply did whatever there was time for. Increasingly, we had to prioritise testing so that we would not slow down development. But quality suffered as a result. We tested only what was absolutely necessary, rather than everything that should have been tested. It became a vicious cycle, and it was frustrating.

On top of that, the tester on a project naturally accumulated the deepest domain knowledge and gradually became indispensable. Everything then depended on that one person and their know how. And that is not a healthy situation, either for the project or for the person involved.

Looking back at the past six months, how has GTS changed the most?

Six months is an incredibly long time at the pace at which we develop GTS. Just consider that a complete rewrite and refactoring of the platform takes our developer only a few days. You could count them on one hand and probably still have a finger or two left. So a lot has happened in that time. We experimented with different models to find those that work best for GTS, fine tuned prompts, implemented a large number of new user requirements, and also worked on the design to make GTS not only easy to use, but good looking as well.

But two changes stand out. First, we have reconsidered and expanded the target group using the application. Originally, GTS was primarily aimed at QA engineers. But we are hearing from more and more clients who want a tool that is simple and more self sufficient. Something that does not require a QA professional at every step.

The second thing, and one I am particularly proud of, is the new chatbot testing module. We had to develop a complete methodology for how to approach the topic in the first place. It is a completely new area where there are no well established paths to follow yet.

Chatbot testing. What does that actually look like in practice?

The module allows you to define a dataset of questions, build an executable test collection, schedule periodic runs, and track statistics and result overviews across all of them. Evaluation is based on the dataset, which defines what the chatbot’s responses must and must not contain.

The greatest value, however, lies in the evaluation itself. We had to develop a methodology that processes the data and interprets the results in a way that is clear to the tester responsible for testing. To make a reliable assessment, you need to work with a large dataset, and that is exactly where AI is ideal.

When you explain to a Project Manager or Product Owner why GTS makes sense, where do you start?

I do not start with numbers. I leave those until the end. I start by understanding and describing the problem. Our goal is not to sell a tool. The reason goes much deeper than that. At the end of the day, we want to leave behind a good product. Quality translated into end user satisfaction is our primary measure. If that works, the numbers are simply confirmation.

Petr Škoda - QA Competence Lead
Do you have a specific example of a project where you were able to measure the impact of GTS?

Yes, and we approached it gradually. We started by taking a smaller part of one project and testing it as a PoC to confirm that our idea would hold up in practice. We then introduced GTS to our QA team and gradually started using it on other projects as well. It soon became clear that the results were even better than we had expected.

For example, test analysis, meaning the creation of test scenarios, used to take around two hours on average. With GTS, it takes 24 minutes. We see similar results in test automation, where the time required has dropped from six hours to 36 minutes.

But the savings are not limited to QA. Thanks to consistent outputs, the traceability of individual artefacts, and the quality of those outputs, the benefits extend across the entire team. We tracked the numbers in a seven person team working in a standard agile development setup. Over the course of a year, we achieved a 13% reduction in total costs.

Is there anything AI still cannot do and where a human is still needed?

We still divide the process into individual phases. It is not a “one shot mode” where I enter a task and then only check the outputs. A human is involved in every phase. That is extremely important to us because we are still responsible for making sure everything is correct and properly reviewed.

We stick to established standards. When we are responsible for quality assurance, we do not interfere with analysis or implementation. We leave that to the people with the relevant expertise. But you still need to address discrepancies between requirements and implementation, as well as ambiguities in interpretation, whether they come from the LLM or from the expert responsible for their particular area.

We do not want to end up in a situation where AI comes up with the requirements, AI implements the code, and AI checks it as well. There still needs to be someone who can hit the brakes and say, “Guys, this just doesn’t make any sense.” Or, “This is so overcomplicated that nobody understands it, and we are generating a lot of unnecessary stuff.” Today, we call that AI slop, and the end user has absolutely no interest in it.

To underline the point, I still believe that a QA engineer can better understand the needs of the end user and make sure the product actually meets their expectations.

There are people behind every AI tool. Who is behind the development of GTS at Green:Code today?

I am glad you asked that, because the success of GTS is primarily down to a great team. Prompt Engineer Marek Novák is the father of the entire concept. AI Developer Jakub Faist then orchestrates the whole solution down to the very last line of code. They are the heart that keeps GTS beating.

Of course, many more people are involved. Our QA colleagues contribute their know how, while our internal design team makes sure GTS is intuitive to use and looks sexy at the same time.

What are you currently working on in GTS, and what can customers expect in the near future?

I have already touched on this a little. We are taking the original idea behind GTS further, and we are approaching it from two directions.

The first is to open GTS up to domain roles as well, rather than limiting it to QA engineers. We want it to be easy to use even for someone who does not work with testing on a daily basis.

The second direction reflects just how different the possibilities are today compared with a year ago. Every QA engineer can now have an agentic system at their fingertips, whether that is Claude Code or our own custom built Green:Code AI orchestration platform. They no longer need as much hand holding.

We have therefore also transformed the core of GTS, which is built around prompts and a highly refined workflow, into standalone agents and skills. This enables more effective collaboration and makes everything move faster. A talented QA engineer can orchestrate the entire process independently and does not need the platform itself to do so.

It is a step that benefits both approaches. GTS retains our know how and standards, while the toolkit gives our QA talent the freedom to work without unnecessary constraints. At the same time, all the benefits mentioned above remain intact.

Finally, how would you briefly introduce GTS to someone who has never heard of it?

If you work in software development and keep hearing that testing capacity needs to be increased, that there is never enough time, or that testing functionality before a production release takes a week or more, GTS will definitely help. Especially if you cannot assess what would genuinely improve the situation because you do not have a sufficient overview of the state of the application.

GTS is not just a tool. It is a guarantee of quality.

No items found.
No items found.
No items found.
No items found.