Getting started is not getting it right

The gap between what AI is reliable for and what it looks reliable for is where damage happens, and enablement fails

4 box matrix reading Not yet, Verify, Experiment, Default to AI
How to lead the change with AI in a nutshell

I’ve been building with AI for three years now. I started by trying to scale UX writing via plug-ins and now have agents, workflows, MCPs, and apps under my belt. I feel comfortable building with AI, and have adopted it into many of my day-to-day tasks.

Yes, I do think AI has gotten better over those 3 years.

But I think we’re telling ourselves a story that isn’t true. AI has gotten very good at one thing:

Starting.

The blank page problem.

Replacing lorem ipsum.

But while starting is faster than ever and error rates may have dropped, the verification burden hasn’t.

I think errors have gotten harder to spot, and models have gotten better at defending them.

For many, getting from zero to a rough first version is the hardest emotional hurdle in any task, and AI clears it in seconds. No more staring at a blank page if you don’t want to.

But I think that’s the ceiling.

Once you’re past “getting started” and into “getting it right,” the story changes.

The hallucination problem isn’t solved.

Yesterday I fed an AI four screenshots of my team’s current budget and asked it to summarize the numbers, analyze them, and combine that with what it knew from past conversations to suggest adjustments for next year. A simple brief coming from someone who knows how to prompt. Verifiable source of truth sitting right there in the screenshots.

It hallucinated two numbers. Not small rounding errors, invented figures. The advice it gave me, built on top of those numbers, was thirty to forty percent off from what would have been logical. I only caught it because I read the output twice and cross-checked the math myself. When I pushed back and said the numbers looked wrong, it didn’t admit the error. It argued with me.

This worries me.

It’s not just that models hallucinate. It’s that they’ll defend the hallucination when challenged, which means the burden of catching the mistake sits entirely on someone who already knows the answer well enough to spot it. If you didn’t know your own budget cold, you’d have walked away with a plan that was off from reality, and you’d have had no reason to doubt it. If you were too busy to pay close attention in that moment, you may have called it a day. For my team, this would have led to real issues later on.

This is just one example. I run into errors and hallucinations in almost every task I prompt, from user research analysis to coding projects. I’ve tried building skills, mds, and MCPs to solve common mistakes with little success. The fact remains: I have to check everything extensively, and sometimes restart several times before the output is workable.

It’s made me paranoid.

A mind map of the different levels of validating AI output
How paranoid are you?

And while the expectations now are that everything should be delivered much quicker because of AI, I feel rushed to ensure the (quickly) generated output holds. And sadly, I don’t think everyone goes through the trouble, leading to more slop than we’re aware of — which we won’t realize until a few months from now.

I tried Claude Fable extensively for two weeks, and while it was noticeably better at first output than Opus, the same pattern showed up: summaries that slipped, details missing, confidence that didn’t match accuracy. (By the way, this is one of the reasons I firmly believe agent tone of voice is a deceptive pattern, as I’ve written about previously.)

Data backs this up: Stanford’s 2026 AI Index found hallucination rates across 26 top models ranging from 22% to 94% on a new accuracy benchmark, and noted that models handle a false statement well when it’s framed as something another person believes, but performance collapses when the same false statement is framed as something the user believes.

It’s not that fast (if you care about quality).

I build MCPs (and with them), work on agent and conversation design, and I code my own website and other small apps despite limited coding experience. The pattern is: I spot an issue, it takes my AI many prompts to find it, and even more to fix it correctly. It’s usually faster for me to open the code and fix it myself. It’s frustrating. And the fixing “minor” errors part eats time I’d rather spend on something else — yes, even getting started from scratch.

So, as someone who’s built a career on establishing functions from 0 to 1 and then scaling them (in my case: content design, localization, and UX), I can say with absolute confidence: AI is not great at scale right now.

The more complex the task, the more sources involved, the more history in the context window, the worse. Especially if you know your domain well (like most senior ICs) and are looking to AI as a way to scale your output, this becomes frustrating and, in my opinion, scary. I keep thinking “would someone with less experience have spotted this? or would they have trusted the AI and built a monster?

I know what the counterargument may sound like: you’re not prompting it right, you’d avoid this with a better setup, better context management, better tooling, yada yada. Sure. But that argument only holds for people who are already comfortable enough with systems and code to build that scaffolding themselves. Most knowledge workers being told AI will make them more efficient are not going to get there.

If a tool only delivers on its promise once you’ve built a sophisticated harness around it, it hasn’t actually solved the problem for the person it was sold to. I’d even argue: it’s not truly accessible then, either.

UX is deteriorating.

This is the part that comes most naturally from my own career. I’ve worked in UX for basically my whole professional life. To me, it’s obvious, immediately, how shallow AI-generated design and copy still are on accessibility, on tone precision, on the kind of judgment a mid-level designer develops through repetition and feedback. Even the newest, most expensive models, explicitly prompted to improve UX and give options, don’t spot what an experienced designer spots instantly.

The 2026 Web for All study measured AI-generated interfaces at 29% compliance across five objectively measurable WCAG criteria, and found that explicitly specifying accessibility requirements in the prompt made compliance worse rather than better.

Screenshot of a card picker frame with low accessibility
A prototype after reading in the design system. Low accessibility across measurable WCAG criteria.

Which means in some cases, using AI for design work costs more than just hiring someone who’s actually skilled at it. Not in dollars necessarily. In rework, in things that ship broken, in the quiet erosion of quality that nobody notices until customers start complaining.

And while AI-generated designs look shockingly similar, I’m continuously surprised at how much models struggle with actual brand and component consistency. Even when they have clear sources of truth to start from, like a sophisticated design system or glossary.

Distortions are adding up.

I don’t fully trust meeting transcripts. I’ve sat in rooms and read the AI summary afterward. It didn’t reflect what was actually discussed. That’s a small example of a much bigger problem: people increasingly making decisions based only on an AI’s summary or reasoning, without ever having been in the room, without ever touching the source material.

The risk is a generation of decisions being made on secondhand, unverified synthesis, by people who have no way of knowing what got dropped or distorted along the way.

This risk is incredibly difficult to estimate right now.

When to and when not to trust AI output is a personal choice and likely relates to people’s personalities (like, their willingness to take risks). I imagine we’ll hear more and more about how different types are using AI and how that affects business outcomes soon.

Additional complexity in Europe.

I’m speaking at HATCH Conference in Berlin next month. My talk keeps circling back to something adjacent to all of this. For years, European product and design teams have imported Silicon Valley team structures more or less wholesale: the org charts, the velocity expectations, the AI-first mandates. That copy-paste approach ignores real differences in culture, compliance, and how people are actually set up to work here versus in the US.

I think it’s undermining AI rollouts across Europe.

Recent research puts US worker adoption of generative AI at around 43 percent, compared to roughly 26 to 36 percent across European countries. A meaningful part of it is structural: European data residency expectations, GDPR-native defaults, and now the EU AI Act, all shape which tools teams can reach for, and how deeply they’re willing to plug AI into real workflows.

This matters for the accuracy question too.

Most AI models are trained and tuned against a default that skews American: American workplace norms, American communication style, American assumptions about what “good” output sounds like. When a European team asks AI to draft something, summarize something, or make a judgment call, it’s not just doing that work against a general knowledge base but also against a cultural default that isn’t theirs. Ask it to draft spend policy copy. The defaults it reaches for are American ones: per diems, tipping, reimbursement norms, no VAT anywhere in sight. Every one of those is a small wrong assumption that someone on a European finance team has to catch and correct. This is a second layer of inaccuracy sitting on top of the hallucination problem.

If we’re serious about AI enablement in European organizations, the local dimension can’t be an afterthought bolted onto a US-built playbook.

Specificity drives change.

I recently went for a walk through beautiful, sunny Stockholm with Matt LeMay, who wrote Product Management in Practice and has been on Lenny’s Podcast talking about impact-first product teams. Somewhere between vintage Italian suits and Pippi Longstocking merch, we ended up comparing notes on what it feels like to work at the forefront of product right now, across our different employers and clients.

The consensus: most companies have not nailed how to actually transform an organization to be “AI first or AI native”, a phrase that gets thrown around constantly and means almost nothing in practice. The leaders who can drive are the ones close enough to building with AI to understand its constraints firsthand.

Most organizations push AI top-down to show employees that it matters, in hopes they will feel forced to adopt it more. But people have already understood it matters, and many have adopted it in their day-to-day work. The reason this isn’t moving the needle for some teams is that there’s no specificity around how to use AI.

What they need is leadership that can draw a clear matrix: here’s what AI is reliably good at, here’s where you can start using it in your day-to-day with confidence, here’s where you can experiment but shouldn’t expect real efficiency gains yet.

Example matrix with clear sections on where AI is reliable vs not
What that matrix might look like

That kind of clarity is what gives people the confidence to adopt something new. Enthusiasm alone doesn’t drive change management.

Be specific.

Where this leaves us.

I think it’s time to stop treating “gets you started” and “gets it right” as the same category of problem. They require completely different levels of trust. Acknowledging this is what can truly move businesses forward with AI, effectively.

We need clear guardrails on what AI is actually reliable for versus where it just looks reliable. Real conversations inside teams about which tasks are safe to hand over completely, which are safe to start with AI and finish by hand, and which are better done from scratch by a person who knows the domain.

The psychological safety teams need to collaborate well with each other extends to working with AI: only if there is trust can transformation and innovation be achieved.

I believe that trust starts with leadership that is honest about the constraints of AI and their expectations.

Nicole is a Content Designer turned Design Director based in Stockholm, Sweden. She potters, writes poetry, and raises little girls in a house by a meadow. You can follow her writing here or get it directly to your inbox via her publication, eggwoman. Nicole is on Linkedin. Her portfolio is nicoletells.com.


Getting started is not getting it right was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.

Need help?

Don't hesitate to reach out to us regarding a project, custom development, or any general inquiries.
We're here to assist you.

Get in touch