← All insights
Automation3 min read

Where AI can genuinely help SAP test design

AI can accelerate document analysis, scenario generation, traceability and failure summarisation — but experienced people must remain responsible for deciding what matters.

Written by
Rufouss Practice Team
Published
2 June 2026
Written for
CIO · Test Director · Engineering Leadership
The point in two minutes
  1. AI is strongest where the task is reading and restructuring large volumes of text — which is most of test design’s preparatory work.
  2. It is weakest exactly where SAP testing is hardest: deciding which business processes and variants matter.
  3. Generated artefacts still need a named reviewer, or the volume becomes the problem rather than the solution.

The question

There is now real AI capability in the SAP testing toolchain rather than only in the marketing. Tricentis announced AI-assisted automated test case generation for SAP in May 2026, including generating end-to-end test cases from natural-language prompts. That is a genuine capability and it is worth being specific about what it changes.

Being specific matters in both directions. Overstating it produces disappointed programmes. Dismissing it means doing by hand work that no longer needs doing by hand.

What we see in programmes

The preparatory work of SAP test design is dominated by reading. Requirements documents, functional specifications, process flows, configuration rationale, prior test repositories, defect histories. On a large programme this is thousands of pages, and it is done under time pressure by people who are also expected to be doing something else.

The consequences of that are predictable and are the root of several problems this publication keeps returning to. Scope gets set from what people remember rather than from what the documents say. Traceability is asserted rather than built. Variants go missing. Not through carelessness — through volume.

This is precisely the shape of problem that current AI handles well.

A better way to think about it

Four places where it earns its keep, all of which share a property: the task is reading or restructuring, and a person checks the output.

Document synthesis. Reading a large specification set and producing a structured summary — the processes described, the variants mentioned, the integration points referenced, the areas where documents disagree with each other. That last one is quietly the most valuable, because contradictions between documents are a reliable indicator of where design intent was never settled.

Candidate scenario generation. Producing a first draft of test scenarios from a process description. The draft will be generic, and it will miss the things that make your landscape yours. It will also produce, in an afternoon, the boring eighty per cent that would otherwise take a fortnight — which frees the experienced person to spend their time on the twenty per cent that requires judgement.

Traceability assistance. Proposing links between requirements, processes and existing test cases across repositories too large to map by hand. Proposals, reviewed. This turns an eighteen-month-old traceability gap from an impossible task into a reviewable one.

Failure summarisation. Clustering the output of a regression run so that forty failures resolve into four underlying causes. This is the least discussed and possibly the most immediately useful, because triage effort is what stops teams running their suites.

The work AI removes is the reading. The work it does not remove is deciding what matters — which is the part that requires having done this before.

What good looks like

A named reviewer for anything generated. Volume is easy now. Two hundred generated scenarios with no reviewer is a worse position than twenty written by hand, because the suite acquires apparent coverage that nobody has validated. The generation is cheap; the review is the control.

Generated artefacts marked as generated. At least until reviewed. This is a small discipline that prevents a large problem later, when somebody needs to know how much of the repository has been looked at by a person.

Claims kept narrow. “We use AI to draft scenarios from specifications, reviewed by a senior consultant” is a claim you can stand behind in a procurement conversation. “AI-driven autonomous testing” is not, and the difference tends to become apparent about four weeks in.

Self-healing understood precisely. It genuinely reduces maintenance for identifier and layout drift — a real and common class of breakage. It does not maintain a scenario whose underlying business process changed, because that is a semantic change. Both facts should be said in the same sentence.

Questions leadership should ask

Which of our test artefacts were generated, and who reviewed them?

What specifically does our tooling do — in a sentence with a verb in it?

Has AI-assisted triage reduced the hours we spend on failure investigation? This is measurable and is the fastest place to see value.

If the tooling were removed tomorrow, which of our claims about coverage would stop being true?

Rufouss perspective

AI can produce a hundred candidate scenarios in an afternoon. Choosing the twelve that matter is still the job, and it is the part that requires having seen this before.

Sources & further reading
  1. SAP Enterprise Continuous Testing by TricentisSAP’s continuous testing offering with Tricentis, positioned around protecting critical business processes.
  2. Tricentis — agentic AI testing for SAP business transformation (May 2026)Announces AI-assisted automated test case generation for SAP, including generating end-to-end test cases from natural-language prompts.
  3. Rufouss field observationPractitioner assessment of where these capabilities help and where they do not.
Written by Rufouss Practice Team SAP Quality Engineering — Rufouss

Written by the senior SAP practitioners who run Rufouss assessment, testing and automation engagements.

Working on something like this?

No question is too simple, and none is too complicated. Ask us what we have seen — there is no charge for a conversation.