Clerk in a utility's market communication back office at two monitors showing a meter photo and an exception queue

AI for exception cases: from pilot to operations

Meter readings, substitute values, master data and rejected messages in market communication and metering

Westnetz clears 90 percent of its meter-reading exceptions automatically. Most German municipal utilities never get past the pilot. It's rarely the model. What's missing is a threshold and someone who answers for it. And a number that proves the benefit.

Summary

An exception case (German: Klärfall) is any transaction in energy market communication or metering that the system can't process on its own, such as an implausible meter reading, a missing value, a master data mismatch between market partners or a rejected message; it affects grid operators, suppliers and metering point operators in Germany and has been under far tighter time pressure since the 24-hour supplier switch went live on 6 June 2025. Exception handling is the AI use case with the most public evidence in the German energy sector: according to a PwC study, Westnetz resolves 90 percent of exceptions in meter reading validation fully automatically, and EWE reads around 50,000 meter readings a month by image recognition with 99.5 percent accuracy. The same study finds 61 percent of utilities stuck in test or pilot mode with AI, and none that consider themselves advanced. Getting to operations takes a confidence threshold with human approval, clear process ownership and three metrics: automation rate, lead time against the process deadline and error rate after release.

Why exception cases are the proven AI use case

A meter reading comes in well above last year's. Billing blocks it, and eventually someone in the back office takes a look. Often enough the reading is fine: the household has put in a heat pump. Checks like this eat a large share of the working day in metering and billing.

Exception case (Klärfall) is what the German energy industry calls any item in market communication, billing or metering that the system can't handle by itself, so a person has to review it: an implausible meter reading, a missing value that needs a substitute, a master data mismatch between market partners, a rejected EDIFACT message.

This is exactly the work that's well documented as automatable. On 13 April 2026 the trade paper ZfK reported on a PwC study of AI in the energy sector. 79 percent of the utilities surveyed have introduced AI solutions, 61 percent are stuck in test or pilot mode, a quarter have an explicit roadmap. Not one calls itself advanced.

79 %
of surveyed utilities have introduced AI (PwC, 2026)
61 %
are stuck in test or pilot mode with it
0 %
consider themselves advanced
90 %
of meter-reading exceptions Westnetz resolves automatically
50,000
readings a month EWE captures by image recognition, 99.5 % correct
~80 %
of manually checked readings turned out to be right all along (Natuvion)

Nearly all the examples the study highlights sit in the back office of grid and metering. Westnetz resolves 90 percent of all exceptions in meter reading validation fully automatically. EWE processes around 50,000 meter readings a month through AI-supported image recognition. In January 2026, EWE NETZ announced it would replace on-site reading with digital self-reading, with an AI checking the photo while the customer is still in the app.

Hand holding a smartphone up to an old black electricity meter in a basement, the register filling the phone screen
Many exceptions start with the customer's photo. Check it in the app and some of them never happen.

In SAP shops this isn't exotic any more either. DSC sells an add-on that pre-scores implausible meter readings directly in SAP transaction EL27 with a machine learning model trained on the utility's own production data. By its own account more than 20 customers use it, among them Stadtwerke Essen, enercity Netz, N-ERGIE and Stadtnetze Münster. Natuvion writes about an ENERGY4U reference project for meter reading validation that has been implemented more than 20 times.

Why here, of all places? Because an exception case has three things AI likes. Volume. A decision with few outcomes: release, correct, ask. And a history in which every earlier decision is already recorded. That's training data most utilities own without calling it that.

Where exceptions come from, and why there are more of them

Regulation drives the load, not AI fashion. Since 6 June 2025 the technical supplier switch in Germany runs within 24 hours, as set by the Bundesnetzagentur in ruling BK6-22-024. In the first week, ZfK reported, only around 10 percent of some 250 grid operators tested could process the new messages correctly, and Stadtwerke München said its number of bilateral exceptions had "almost increased tenfold" (our translation). An exception that could once sit for a few days now holds up a switch that's meant to be done in 24 hours. Details are in our article on the 24-hour supplier switch.

1 October 2026 brings the next push. With notice no. 56, 25 AHB and MIG versions for 13 EDIFACT message types become binding, along with AS4 profile 1.2. New validation rules mean new rejection reasons, and every rejection is somebody's exception. What goes live, exactly, is covered in our piece on notice no. 56. From 1 April 2028 the migration to the MaBiS hub starts, and grid operators will mainly be supplying master data.

Typical exception cases, who has them and what a model can contribute
ExceptionTypical roleWhat a model can do
Implausible meter readingGrid operator, metering point operator, billingScore the value against history, weather and asset data, propose release
Missing value, substitute valueMetering point operator, grid operatorEstimate the substitute and report how good it is
Customer meter photoMetering point operator, customer serviceRead meter number and register, reject bad photos on the spot
Master data mismatchGrid operator, supplierFind duplicates and the likely link between market and metering location
Rejected message (APERAK)All market rolesClassify the rejection reason and route it to the right desk

We've kept the table sober on purpose. No model repairs a market location that's set up wrong in your own system. It just finds it faster. Why clean master data comes first is explained in our article on master data quality in market communication. For feed-in customers, the GPKE and MPES processes add their own cases, described in GPKE and MPES for feed-in processes.

Why the pilot never reaches operations

PwC also asked about obstacles. 30 percent point to missing technology, 28 percent to a lack of skilled staff, 18 percent to legal or ethical concerns and 16 percent to resistance inside the organisation. The BDEW study Digital@EVU 2026, published on 14 April, gets to the same place from another angle: a third of utilities have implemented their own AI strategy, and many of the core hurdles identified in 2023 are still there.

We see four causes hiding in those percentages, though they're rarely named that way.

Data quality comes first. A model that learns from three years of exception decisions learns their mistakes too. If a mismatched metering location got waved through by hand for years, the model treats it as correct.

Then ownership. In the pilot the project team decides. In operations somebody has to sign off that an automatically released reading may go onto a bill. Billing, metering or IT? Until that's settled, every release is only a recommendation, and the back office checks twice.

Market communication has its own rhythm, too. A model trained in spring on the old error patterns meets new message versions on 1 October. Without a retraining plan it gets worse on the very day the exceptions pile up.

And the operating model. Who watches the hit rate, who changes the threshold, who switches it off? In the grid portfolio we've run since 2024 for a multi-utility in northern Germany, around 40 initiatives through 2030 compete for the same people. An AI pilot without an owner sits next to projects that have a statutory deadline. You can guess how that ends.

Approval or autopilot?

How much should a model decide alone? Two vendors that both sell exception-handling AI answer differently.

Natuvion, on 5 May 2026, argues from the cost of today's checks. About manually reviewed meter readings it says:

In around 80 % of cases, the check only confirms that the meter reading was correct.

Natuvion, Christian Pfister (our translation),

Natuvion calls that "a lot of effort for a simple release" (our translation). The conclusion seems obvious: let the model release plausible values itself and send only the doubtful ones to people.

Pexon Consulting pushes back on 7 September 2026, on the neighbouring case of implausible bills: "Implausible bills remain an approval case, not an autopilot" (our translation). The model may flag the amount and the meter, it says, but it may not cancel anything. Westnetz's 90 percent is the state of the industry, not a promise anyone can copy into a service level.

And BDEW reminded everyone in April 2026 that "many of the core hurdles identified back then still exist" (our translation), three years after its previous survey.

Who's right? They're describing different failures. With 80 percent of values correct, manual review mostly costs time. With a cancelled bill, a mistake costs the customer's trust. We won't settle this across the board. The line has to be drawn per exception type, and it has to be written down.

Human in the loop, and the benefit in numbers

In exception handling, "human in the loop" isn't a slogan. It's a number: the confidence threshold. Above it, the system releases and logs. Below it, a clerk gets a proposal with reasons. And a small sample of automatic releases still crosses someone's desk, so you notice when the threshold stops fitting.

Flow: exception case, check by rules and model, automatic release or decision by a clerk, feedback into training
Rules and a model check every exception. Confident cases are released automatically, uncertain ones go to a person, and every decision feeds back.

No study tells you where the threshold belongs. It depends on what a mistake costs. A wrongly released meter reading ends up on the annual bill and comes back as a complaint. A misrouted APERAK costs an hour of searching.

So we measure the benefit with three metrics, and we take them before launch, not afterwards.

Rate
How many exceptions close without a human touching them
Time
How long a case stays open, measured against the process deadline
Errors
How often a release has to be corrected later

That third number is missing from almost every success story. An automation rate of 90 percent means little until someone knows how many automatic releases come back as complaints. Without a baseline from before the model, you can't defend any of the three, not even to your own controlling team.

Two market communication staff at a whiteboard with exception notes such as check substitute value and human approval
Which exceptions the model may close alone is for the team that handles them today to decide.

The back office's decisions are more than a control. They're next quarter's training data. Capture them with reason and outcome, and after the October format change you'll have a fitting model again within weeks.

EU AI Act and MsbG: what applies to exception-handling AI

Good news first. The European Commission's draft guidelines of 19 May 2026 on classifying high-risk systems name one metering case explicitly as not high-risk: an AI system that analyses images of installed electricity meters and gives the fitter recommendations on errors and improvements. It has no direct safety function. The draft isn't final. Still, exception-handling AI that checks meter readings or sorts messages sits even further from a safety component under Annex III point 2 than that example. How the other utility cases fall is covered in our article on Annex III and the Commission's guidelines.

It's different once the AI talks to customers. The transparency obligations under Article 50 have applied since 2 August 2026. A bot that explains to a customer why their reading wasn't accepted has to be recognisable as AI. What that looks like in customer service is in our piece on AI in utility customer service.

On data protection, Germany's Metering Point Operation Act (MsbG) is stricter than the GDPR alone. Section 50 lists the permitted purposes, billing and balancing among them. Section 66 lets the grid operator process meter values only where strictly necessary, for grid usage billing for example, and paragraph 3 requires personal meter values to be deleted or anonymised after fixed periods. Section 70 allows processing beyond that only without personal reference. For an exception model that means two things. The training history needs a deletion concept before it grows. And a model trained on grid data for billing doesn't simply move over to sales; why not is explained in our article on meter data, AI and unbundling in the multi-utility. This is general guidance, not a review of your specific case.

Where it can go wrong

Automatic releases make errors quiet. A person who waves through a wrong value gets noticed sooner or later. A model that waves through a thousand values a day gets noticed at the annual bill, a thousand times over.

Bought-in tools add a second catch. If the rules, thresholds and training data live with the vendor, a utility that switches vendors loses its exception knowledge along with the contract. Who owns the decision history belongs in the tender.

An automation rate without an error rate after release isn't a success story. It's a bet on the next wave of complaints.

What utilities should do now

October is a good moment, because it brings new exceptions anyway. Measure now and you'll have something to compare against afterwards.

Five steps from pilot to operations

  1. Count exceptions by type

    For four weeks, record volume, handling time and outcome for each exception type. Most utilities know the total but not the split, and the split decides where a model pays off. Note how often a case runs up against a deadline, too.

  2. Pick one type

    Start with the one that's frequent and cheap to get wrong.

  3. Write down who owns it

    Decide who is accountable for automatic releases, who may change the threshold and who switches it off. It's one sentence in a process description, and it's the most common gap between pilot and operations.

  4. Set threshold and sample

    Start with a high threshold and a large sample. Lower both only once the error rate after release has held steady for several months. Build retraining into every market communication format change, the next one on 1 October 2026, then at the Bundesnetzagentur's pace.

  5. Document law and data

    Purpose under the MsbG, deletion concept for the history, an entry in your internal AI register, Article 50 labelling if customers see the answer.

How we set up projects like this, from a prototype on your own exception data to governance for day-to-day operations, is described on our page AI for utilities. For AI on the forecasting side of the same value chain, see our article on AI in balancing group management.

Further reading

Frequently asked questions

An exception case (German: Klärfall) is an item in market communication, billing or metering that the system can't process automatically, for example an implausible meter reading, a missing value, a master data mismatch between market partners or a rejected EDIFACT message. A person has to review and decide. Since Germany's 24-hour supplier switch went live on 6 June 2025, many of these cases run against much tighter deadlines.

Meter reading validation has the best evidence: according to PwC, Westnetz resolves 90 percent of these exceptions fully automatically, and EWE reads around 50,000 meter readings a month by image recognition. Substitute values, matching master data mismatches and sorting rejected messages by reason also fit well. What matters is high volume, few possible decisions and a documented history.

According to the PwC study, 61 percent of utilities are in test or pilot mode with AI. They cite missing technology (30 percent), a shortage of skilled staff (28 percent), legal concerns (18 percent) and organisational resistance (16 percent). In practice what's often missing is clear ownership of automatic releases, a retraining plan for format changes and a baseline measured beforehand.

Probably not, going by the European Commission's draft guidelines of 19 May 2026. The draft explicitly names image checks of installed electricity meters as not high-risk because they have no direct safety function. The draft isn't final. If an AI talks to customers directly, the transparency obligations under Article 50 have applied since 2 August 2026.

Section 66 of the German Metering Point Operation Act (MsbG) lets grid operators process meter values where strictly necessary for listed purposes, such as grid usage billing, and requires personal values to be deleted or anonymised after fixed periods. Section 70 allows processing beyond that only without personal reference. An exception model therefore needs a deletion concept for its training history. Whether a specific use is permitted is for the legal department to assess case by case.

With three metrics taken before launch: automation rate, lead time measured against the process deadline, and error rate after release, meaning how often an automatically released value has to be corrected later. An automation rate on its own says little.