← All work

Case study

SAP HANA regression readiness

A repeatable, AI-assisted readiness workflow for SAP HANA analytics changes: automated multi-phase regression checks, specialist-reviewed findings, and controlled reporting — with production activation cadence rising during its staged adoption.A repeatable safety check for changes to a big corporate reporting system. The checking is automated and always runs the same way; a human specialist still reviews anything it flags before a word reaches a stakeholder, and every run leaves a dated record someone else could repeat.

Maturity: Operational
MaturityOperational
TopicsAgent orchestration · Data & analytics · Human governance
TechnologiesSAP HANA · AI agent skills · Model-tiered LLM routing
Reading levelSame figures at both levels.

Problem

In an enterprise analytics and reporting environment, a recurring question precedes every promotion: has this change to a SAP HANA analytics object introduced a regression, schema breakage, or a data anomaly? Performed manually, that validation varies between analysts and is difficult to reproduce. Ambiguous findings — unusual measure deltas, schema drift, distribution shifts — are not investigated consistently, and stakeholder reports rarely distinguish findings that align with the planned change from findings that do not.

Before any change to a big corporate reporting system goes live, someone has to answer the same question: did this change quietly break something, or make the numbers wrong? Done by hand, the answer depends on which analyst does the checking and is hard for anyone else to repeat. Odd results get chased inconsistently, and the report that finally reaches the business rarely separates the differences that were expected from the ones nobody planned.

The manual baseline was concrete: performed by hand with spreadsheets and queries, a validation pass took days to a week — and still covered only a fraction of the surface, spot-checking three measures on the principal view with no dimension-level quality checks at all.

The old way was slow and thin: working through spreadsheets and hand-written queries, one pass took anywhere from days to a week, and even then it only sampled three of the numbers on the main report and checked none of the categories they were broken down by.

Constraints

In plain terms first: the automation is never allowed to publish its own conclusions, the part that writes the report cannot change what the report says, and raw figures never travel into a shared document. The precise ground rules:

  • No automated finding reaches a shareable report without specialist review; the review step deliberately excludes automation.
  • Analysis and report formatting are separated: the report-writing component cannot re-examine source data or alter conclusions.
  • Raw data figures stay in the technical output, never in the shareable report.
  • Every run must leave a dated, reproducible artifact with per-finding reproduction steps.
  • This page identifies no employer or client, names no internal objects or tables, and publishes its one metric as an association rather than a causal result.

Architecture

The core is a reusable validation skill that executes a fixed six-phase regression sequence — schema contract, label quality, measure totals delta, attribute distribution, null coverage, and unit-level sampling — in the same order for every object and release:

At the centre is a reusable checking routine that always runs the same six steps in the same order, whatever is being changed: does the structure still match what was promised, are the labels sound, have the totals moved, have the category breakdowns shifted, are there gaps where there should be values, and do individual records still look right when sampled:

  • Two modes — a two-environment comparison for promotions, and a single-environment baseline.
  • Structured findings — every phase emits severity classification and reproduction steps.

Around the skill sit three deliberately separated components:

Around that routine sit three parts, kept deliberately apart so that no one of them can do another’s job:

  • Analysis agents — run the phases.
  • A structured triage reference — guides human investigation of flagged findings.
  • A report-writing subagent — assigned to a lower-cost model tier; it formats established findings for the audience and cannot perform analysis or change a conclusion. The separation is enforced by role definition and model assignment.
Readiness workflow from planned change to retained run evidence A planned change enters readiness assessment and automated regression checks run six phases producing structured findings. A specialist reviews flagged findings with LLM assistance. If ambiguity remains, a reinvestigation loop applies structured triage and returns to review. Once findings are explained, a lower-cost report-writing subagent produces the controlled report, and dated reproducible run evidence is retained. Planned change enters readiness assessment Automated checks six phases · severity · repro steps Specialist review LLM-assisted · triage-guided Reinvestigation discriminator-first triage Controlled report lower-cost subagent · DOCX Run evidence retained dated · reproducible ambiguity findings explained
Recreated diagram of the readiness workflow. Names and environments are withheld.

AI and agent workflow

A specialist opens a readiness assessment for the target object, naming the source and target environments and the review window. Then:

A specialist starts a readiness check on whatever is being changed, naming where it is coming from, where it is going, and the period to look at. Then:

  • The validation skill runs the six phases and returns the structured findings table.
  • The specialist reviews every flagged finding with LLM assistance, guided by a discriminator-first triage reference — classification against the planned change’s scope is explicitly plausibility-based, never causal confirmation.
  • A decision gate follows — findings that cannot be explained by transaction composition, distribution shift, or planned-change scope go back into reinvestigation before any report exists.
  • Only accepted findings reach the report-writing subagent, which produces the audience-framed DOCX deliverable.

The deeper difference from the manual stack appears after a finding. With the specialized agent and skill, a flagged delta can be interrogated in place — drill into the contributing dimension, isolate a slice, re-run a targeted phase — without leaving the governed session. The spreadsheet-and-query workflow this replaced dead-ends at exactly that point: beyond authoring new queries or building pivots from scratch, there is no follow-up path.

The bigger change from the old way shows up after something is flagged. Now the odd number can be chased straight away, in the same session — break it down, narrow it to one slice, re-run just that check — without stopping to build anything. The old spreadsheet-and-query approach stopped dead at exactly that moment: the only way onward was to write fresh queries or build a pivot table from scratch.

Human governance

Where a person is required, and what that person actually does:

  • High, Critical, and Breaking findings never pass directly to report generation — a specialist reviews them first, and the triage protocol explicitly excludes automation from that step.
  • The review is real rather than ceremonial — run records carry human-applied corrections with explicit correction notes, showing the specialist changed automated output before a record was considered final.
  • The report writer cannot alter a conclusion — what the specialist accepted is what the stakeholder reads.

Evidence

Each run produces a dated, self-contained artifact:

Every run leaves behind a dated record that stands on its own:

  • All phase outputs and the flagged-findings table, with severity and reproduction steps.
  • A findings confidence score and any human correction notes.
  • Runs are therefore re-examinable and comparable across releases; the activation-date series behind the metric below was derived reproducibly from platform metadata.

Equally important is what this evidence does not establish — and what this page therefore does not claim: causation, defect reduction, testing-time savings, cost savings, or adoption beyond the observed runs.

One class of statement on this page rests on a different footing: the manual-baseline comparisons — cycle effort, coverage, and the drill-down dead end — are the operator’s direct experience of the prior process. They are attested rather than artifact-derived, and labelled here as such.

Outcome

Established a standardized, repeatable regression-readiness workflow that produces reproducible run evidence and controlled reporting artifacts for planned SAP HANA analytics changes, with human-reviewed findings and an explicit reinvestigation gate before any shareable report is produced.

Delivered a consistent, repeatable safety check for planned changes to the reporting platform. Every run leaves evidence someone else could reproduce, every finding is reviewed by a person, and anything unexplained has to go back for another look before a shareable report exists at all.

Coverage also widened: where the manual pass spot-checked three measures and no dimensions, each automated run now validates every measure on the view and quality-checks every dimension.

How much gets checked also grew: the manual pass sampled three numbers and none of the categories behind them; each automated run now checks every number on the report and every category it can be broken down by.

Across the observed production metadata for the principal analytics object:

Measure Earlier period Later period
Average interval between production activations ≈ 44.5 days ≈ 25.4 days

That is roughly a 43% shorter activation interval, or about 75% higher activation frequency. The readiness practice matured in stages across that window — manual assisted checks, then IDE-integrated AI support, then the first packaged agent-skill runs — so the comparison reflects the practice’s maturation period. The figures demonstrate increased change throughput during that period; they are published as an association, not as proof that the workflow was the sole cause or that defect rates declined.

Lessons

What I would carry into the next system of this kind:

  • Stating the boundary of a metric is itself a professional differentiator: “association during the maturity period” survives scrutiny that a causal claim would not.
  • Separating analysis from report formatting keeps conclusions stable and makes model tiering natural — the expensive reasoning happens where it matters, and the formatting layer cannot quietly change a finding.
  • A discriminator-first triage sequence turns “this delta looks odd” into a bounded investigation instead of an open-ended one.
  • The step-change is not speed alone but diagnosability: an agent that can follow a finding down beats a report that can only restate it.
  • A finding without reproduction steps cannot be governed; reproducibility is what makes the human gate meaningful.