Celonis · AI · Enterprise · Product Design
June – September 2024 · GA November 2024Annotation Builder
- Role
- Lead designer, frontier AI products
- Team
- Product Manager · Frontend squad (3) · Infrastructure engineering
- Could kill it
- VP of Engineering & Product
In June 2024, Celonis's executive team committed to unveiling the company's AI offering at Celosphere — the annual keynote, four thousand customers in the room, October.
At that point the platform had no AI product. It had a chatbot.
Twelve weeks, and a stage that was already sold.
The knot
Everything that made this hard was upstream of the design
The problem itself was legible enough. Celonis customers run process mining over operational data — supply chain, order management, manufacturing — and a large share of that data is free text. Finding duplicates, errors and divergences across hundreds of thousands of rows is work that takes teams of analysts weeks. Sales wanted it solved. Our own validation said customers wanted it solved. On the what, we were aligned from day one.
The difficulty was everywhere else.
Validation had to fit inside the build
Twelve weeks from kickoff to ship meant research, design, testing and handoff overlapped rather than sequenced. There was no version of this where we explored broadly and then chose. We had to be right early.
The design system set the ceiling
Whatever we designed had to be buildable from the components and the Angular library we already had. There was no room to design new interactions and contribute them back — so the dependency ran backwards. I had to pull the design system team in to build primitives for us: prompting surfaces, AI disclaimers, a way to mark data as AI-enriched.
Nobody knew what prompting UI looked like yet
This was 2024. How long should a prompt be? How much context is enough before it becomes confusing? How do you expose retrieval or web search in an interface? There were no established patterns to borrow — internally or in the market.
We couldn't write back to the data model
Writebacks to the Object-Centric Data Model weren't possible yet. We could detect and modify divergences in a dataset, but couldn't return the result to the model. That capability depended on other infrastructure teams and only landed as a core feature in March 2026 — well over a year after launch.
My angle
A knife, not a swiss army knife
Process Copilot was becoming a swiss army knife where none of the tools worked properly. I wanted to build one knife that cut exceptionally well.
That was the argument, and I made it with the PM against most of the room. It carried a second commitment that turned out to be harder to defend: the configuration had to stay radically simple. No tool configuration. No advanced mode. No "pro users will figure it out."
Give the data. Prompt the job. Say what you expect back. Three steps, and nothing else on screen.
The competitive study is what made us confident. Palantir and SAP had both shipped solutions to roughly this problem, and both had answered it with dense configuration surfaces — input fields, checkboxes, toggles, panels — for a task that in plain language is "read this text and tell me what it says." We set ease of use as the thing that would differentiate us, and then made it an exit criterion rather than an aspiration.
We won the argument with prototypes, not decks. Vibe coding was new in 2024, and putting something working in front of the CTO, the executives and the AI task force created a reaction that no specification could. Once the AI task force and go-to-market validated that this was the fastest path to something excellent, the buy-in followed.
Without itWithout that fight, what shipped would have been another enterprise configuration dashboard — input-heavy, option-rich, and silent on the only question that mattered: what is the user actually trying to do?
The fork
Three ways to build it. Two of them died.
The interesting thing about these two dead ends is that they died of opposite causes. The first had nothing wrong with it except that it was ordinary. The second was the best work of the three, and the stack couldn't build it.
The case for it
It was what users described when we asked them. Engineering was comfortable with it, because it was assembled entirely from components we already had. It would have shipped on time without drama, and it got most of the attention in the first round of reviews.
How far it got
Figma, then a vibe-coded prototype covering every capability
What killed it — Product — not Engineering
No test killed this. We did. There was no WOW in it. The experience was blunt and generic, and what we'd have shipped was a competitor's product with our logo on it.
What it taught
People wanted the capability badly enough to tolerate any interface wrapped around it. That meant the interface was entirely ours to decide — and once it was ours to decide, "tolerable" stopped being an acceptable target.
Artifact not recoveredThe case for it
It had everything the first path lacked. The prototype landed with real force, and its mental model matched the actual output — you were building the thing you were going to get, rather than filling in a form that would eventually produce it.
How far it got
Prototype, tested with users
What killed it
The stack couldn't build it. Emotion — the Angular library behind the Celonis UI — had no primitives for a node canvas, and the development pipelines couldn't be aligned inside the timeline. Shipping it meant either missing the keynote or taking on significant custom-interface debt. That's what decided it.
What it taught
Testing separately showed the node canvas wasn't landing with the data engineer persona — they couldn't grasp the interface — which made the constraint easier to accept. But the honest sequence is that the stack decided it, and the research confirmed it afterwards.
Artifact not recoveredThe case for it
It was buildable from primitives the design system already had, which meant it could ship inside the timeline without debt. And it tested clean — cleaner than either alternative.
How far it got
Shipped
The path that shipped: data in, prompt, declared output.
The second path isn't fully dead. Its core idea — that you should be looking at the output while you build it, not imagining it — survived into the shipped product as the live data preview. The canvas was wrong. The instinct behind it wasn't.
The turn
The signal was a question nobody asked
In session after session, people asked how we'd extend it and what was coming next. Nobody asked how to do the thing in front of them.
We were testing with the AI task force and value engineering — internal stakeholders close enough to the customer problem to be a fair proxy. I went in watching for the usual things: hesitation, misclicks, the moment someone's hand stops moving.
What I got instead was people skipping ahead. They wanted to know what else it could categorise, whether it would handle their edge case, what the roadmap looked like. The configuration was straightforward enough that it never became a topic. Not one "how am I going to do this?"
That absence is what told us the experience had landed. People who are stuck ask about the present. People who aren't ask about the future.
Why it worked
The bet on progressive disclosure was simple: users do one task at a time. Prompt composition was the genuinely ambiguous step — it took people minutes, and it was where the whole thing succeeded or failed — so I wanted it isolated, with nothing else competing for attention while they worked on it.
The prompt assistant did the rest. Give it a one-line description of your goal and it composes a full prompt for that case. In the short term it removes the blank-page anxiety, the "what am I supposed to type here." Over time it does something more useful: users read the expansion, and learn what a good prompt actually looks like.
The craft
Three decisions I'd defend again
The output is visible from the first click
Whatever you select as data input renders immediately as a preview table. The output isn't described, it's shown — so no step in the pipeline is ambiguous about what it's contributing.
Evidence and taste
The prompt assistant, not a second chatbot
A one-line goal in, an elaborated prompt out, filled directly into the prompt field where the user can edit it.
What else was on the tableWe considered routing this through the existing chatbot. Two AI interfaces side by side, overlapping in purpose, was too much interface for one screen — so it became a function of the prompt step instead of a companion to it.
Evidence and taste
Splitting the output out of the prompt
The output column gets its own step: pick its type, give it a name. It's no longer something you describe inside the prompt and hope for.
What else was on the tableThe other direction was splitting the prompt further still — separate fields for role, expectations, examples. That made the experience worse and the backend harder, so we stopped at one split.
What the testing showedTesting showed that when role and expectations lived in a single prompt, people wrote vague output descriptions and got vague results back. Funnelling the output into its own step removed the guesswork — no house style to learn, no getting used to it.
Decided on evidence
Invented, because nothing existed
- The prompting interfaceThe prompt surface and its interaction patterns were Celonis's first. No internal precedent existed, so the patterns were designed from scratch and later absorbed into the design system.
- @-syntax for data referencesTyping @ inside a prompt surfaces the available data and inserts it as a chip, so a prompt can reference specific fields precisely rather than describing them in prose.
- Retrieval and web search as UIHow RAG and web search should appear, be enabled and be explained in an enterprise interface — all of it new, all of it now part of the system.
Shipped knowing better
- The patterns shipped genericThere was no time to iterate on taste during development. We went back the following year with the design system team and built the AI library that's now part of the system — but what launched was rougher than it should have been.
- Inline prompt improvement was cutA Grammarly-style pass that highlights a sentence in your prompt and shows how it could be sharper. The development cost didn't fit the deadline.
The proof
What it did, and what it cost
Annotation Builder was presented at Celosphere in October 2024 as the company's main AI offering, and reached general availability that November. The early-access sheet filled with more than three hundred customers asking to be in the public preview group.
- Kickoff to ship
- 12 weeks
- June 2024 kickoff, September 2024 ship
- Customers rolled out
- 338
- Internal tracking · Since GA, November 2024
- How well it met expectations
- 4.08 / 5
- Post-release customer survey · November 2024 · n = 39
- ARR influenced
- $230M+
- Internal tracking · FY2025
- Productive customers from launch
- 30+
- Internal tracking
The configuration is well structured and easy to implement. Ticket routing took only one day from development to production.
Telekom
AI is not perfect yet, but Annotation Builder works great for us. Free-text categorization is especially promising.
Dell
It is a powerful tool that performs exceptionally well. The configuration process is straightforward, and the results align with expectations. During the AI Labs, it took one day to reach productivity.
Nestlé
I am very positively surprised by how Annotation Builder solves complex tasks through reasoning. It aligns with our customer problems.
WEIR Group
I was surprised by how well it captures fuzzy business rules and produces well-defined decisions. Prompt-based configuration makes it easy to build.
ThyssenKrupp
What it left behind
- The AI group stopped existingCelonis became an AI-first platform in the months after launch, and the separate AI group was dissolved into the rest of the product organisation. The argument against a single AI surface had been won in a way that outlasted the feature.
- The patterns became the systemThe prompting UI, the @-syntax and the AI disclaimers built for this project became the foundation of the AI library the design system team and I built the following year.
- It changed how I workPrototyping in code collapsed my validation cycles and made me a contributor inside the software development lifecycle rather than an input to it. It also made me the person the team comes to on AI products — which is a different job than the one I had in May.
What it cost
The finesse. We dropped the UI craft along the way — the microinteractions, the considered colour, the detail that makes an interface feel made rather than assembled. What shipped was a typical Celonis interface, and I knew it while we were shipping it.
Same brief, second attempt
Push the frontend harder on craft. Same twelve weeks, same scope — but I'd spend more of my own capital on the microinteractions and the colour implementation instead of banking all of it on the structural argument.