AI Reality Check: Inside Two Manufacturing Enterprise Rollouts
Two client case studies — a furniture fabric manufacturer and an automotive filter manufacturer — on messy data, ERPs, first pilots, and the road from a working prototype to a solution the whole business can trust.
Key Takeaways
✅ The problem is usually the data, not the AI. Fix your source of truth before you touch the model.
✅ A CEO's weekend prototype proves demand, not readiness.
✅ Trust comes from process. Users don't care how impressive a demo is — they care if it holds up tomorrow.
✅ Every new use case resets the risk. Answering "where's my order?" and canceling that order are not the same level of trust.
What happens when you ask AI to make sense of a business with 125,000 products, over 1,000 files, and years of documents that contradict each other? And what happens if you skip the "ask" part entirely — connect Claude to your ERP over a weekend and just put a solution to work?
We tested both approaches in practice. In the first case, a furniture fabric manufacturer with a staff of nearly 1,000 handed us everything they had: catalogs, internal documents, spreadsheets, and years of accumulated information. No clear structure, no instructions on where to start. Essentially: "here's our data — figure it out."
In the second, an automotive filter manufacturer with around 700 employees. The CEO didn't wait for an outside team. He opened Claude himself, connected it to the ERP — and within a single weekend had a solution that actually worked. But only up to a point.
In this piece, we'll cover what happens after the initial "wow" moment fades. Where AI genuinely saves time. Why a solution built solo can quickly hit a ceiling. And what it takes to turn a pilot into a tool the whole business can rely on.
We'll walk through two real cases — Kravet and S&B Filters — and answer:
— why low accuracy at the start can actually be a good result;
— why AI quality depends primarily on data and how it's integrated into the workflow;
— why a CEO-built prototype shouldn't be dismissed — or scaled as-is;
— how to tell whether a company is genuinely ready for enterprise AI.
Kravet: The Assistant That Got 4 Out of 10 Answers Wrong
Kravet Inc. brings together the Kravet, Lee Jofa, Brunschwig & Fils, GP & J Baker, and Donghia brands. The company has been operating since the 18th century, employs nearly 1,000 people across multiple countries, and its product archive covers more than 125,000 SKUs.
For a fabric manufacturer, that scale doesn't just mean a wide range of products. It means decades of accumulated information: product specs, documents, spreadsheets, sales materials, operational rules, and customer-facing content.
You'd think this is exactly where AI should deliver the biggest impact. If an employee can ask one question instead of opening several systems and hunting for the right file, the company saves time across every department.
But first, you have to figure out where the correct answer actually lives.
Insight 1: The Problem Wasn't the AI — It Was That the Company Had No Single Source of Truth
Sales, procurement, operations, and HR teams were all working off one massive but highly inconsistent pool of data: PDFs, spreadsheets, scanned documents, blog posts, and a product catalog in Algolia.
Some documents were outdated, duplicated other information, or flatly contradicted each other — even when describing the same product. SKU search wasn't reliable either.
For a person, that's already a tough situation. But someone who's been at the company for a while carries context: which document is newer, what the product is called internally, who to ask when something's unclear.
AI has none of that context. It sees the sources and tries to build an answer from what's available. When sources conflict, the system can still produce a confident-sounding answer — and that's exactly what makes the error dangerous.
So AI doesn't clean up messy data automatically. If anything, it makes the mess more visible. The system quickly surfaces contradictions that used to be scattered across dozens of manual processes.
Kravet had already tested off-the-shelf AI tools and concluded they couldn't handle a dataset like this. The team didn't argue with that conclusion or launch a large-scale build blind.
AI doesn't clean up messy data automatically. If anything, it makes the mess more visible.
Insight 2: A Pilot Should Expose Weak Spots, Not Create an Illusion of Readiness
BotsCrew ran a three-week, low-risk pilot. The assistant was trained on Kravet's data exactly as it existed at the time. The result was far from ideal: accuracy came in below 60%. Roughly 4 out of 10 answers were wrong or incomplete.
At this point, it would've been easy to say: "AI just doesn't work for us." But that would've been the wrong conclusion. The pilot's job wasn't to prove the system was ready to launch — it was to show what was standing in its way.
Testing surfaced five distinct problems:
— conflicting sources triggered confident-but-wrong answers;
— scanned PDFs weren't being read properly and effectively dropped out of search;
— product pages weren't indexed, breaking SKU search;
— the system was inconsistent about which sources it picked, so the same question could yield different answers;
— test questions were clustered around a few topics — which, as a side benefit, revealed which scenarios mattered most for the future product.
Low accuracy turned out not to be a final grade, but a roadmap.
Insight 3: What Often Improves AI Quality Isn't a New Model — It's Discipline Around It
None of the five problems required switching to a "smarter" language model. What needed to change was how the system handled information:
— strip out outdated and duplicate files at the source level, rather than trying to correct for them on every query;
— move to a 128,000-token context window so answers could draw on more material;
— pull in more sources per answer and cross-check them, instead of trusting the first match;
— lower the model's temperature, trading some creativity for more consistent results.
After that, accuracy rose from under 60% to nearly 90%. This case matters for another reason too: it challenges a common assumption about AI development. Public discourse often makes it seem like the key decision is picking the right model. But in an enterprise setting, the model is just one part of the product. Sometimes a company spends its time hunting for a new LLM when the real problem is an old PDF, a duplicate file, or the lack of a defined source of truth.
Sometimes a company spends its time hunting for a new LLM when the real problem is an old PDF, a duplicate file, or the lack of a defined source of truth.
Insight 4: Trust in AI Is Built Through Process, Not a Single Good Result
Kravet's CTO specifically called out the process of working through failed answers. The team didn't try to rush past problems to show off a polished demo. They dug into the errors and used them to drive the next iteration.
For enterprise AI, this matters a great deal. Users don't judge a system by how impressive one answer is. They judge it by whether they can rely on it tomorrow, next week, and in an unfamiliar situation.
That's why trust doesn't come from a promise of "nearly 90% accuracy" — it comes from understanding where that number came from, how it's being checked, and what happens to the rest of the answers.
Insight 5: The Best Proof of Value Isn't a Post-Pilot Score, but a Request for New Use Cases
Kravet didn't stop at the internal assistant. Once trust in the first solution was established, two more projects emerged. The first was a Sales Agent, migrated from an old BI stack to BigQuery so sales managers could get sales and inventory information in natural language instead of relying only on static dashboards.
The second was a Shopping Assistant for customers, integrated into Kravet's storefront on Adobe Commerce and Algolia search. The assistant combines product data with marketing copy, asks clarifying questions when a customer's requirements conflict, and supports cart functionality for logged-in users.
That's what trust in enterprise AI actually looks like — not an enthusiastic reaction to a demo, but a business's willingness to hand the system harder work.
S&B Filters: When the CEO Built the First Prototype Himself
S&B Filters manufactures high-performance automotive air filters in the U.S. The company employs over 700 people, and its sales, order fulfillment, and customer support all run on NetSuite.
Insight 6: The Best Use Case Is Often Hiding in the Most Boring Process
"Where's my order?" rarely sounds like a strategic question. But for a support team, it can become the single biggest source of workload. Every request meant a manual lookup in NetSuite. As inquiries started coming in by phone, email, and Shopify, the same process repeated itself dozens or hundreds of times.
For the customer, it looked like waiting for a response. For the agent, it looked like constantly switching between the message and the ERP. For the business, it looked like a bottleneck that grew right alongside order volume.
In processes like this, AI doesn't need to be "innovative." Its value can simply be removing a repetitive manual action that's no longer adding value.
AI doesn't need to be "innovative." Its value can simply be removing a repetitive manual action that's no longer adding value.
Insight 7: A CEO-Built Prototype Is Proof of Demand — Not a Finished Product
S&B Filters CEO Barry Carter didn't wait for an outside team. He connected Claude directly to NetSuite via MCP, wrote his own prompts, and built an internal assistant for checking order status.
The prototype did its main job: it proved the problem was real and solvable. But it wasn't ready for the whole team to use. Response time was 4–6 minutes. Handing the tool to more than one person risked it buckling under load.
The key here is avoiding two extremes: dismissing the prototype as "not real AI" because the CEO built it over a weekend, or deciding it was already scalable as-is. The more useful way to look at it is as a highly valuable signal. The company hadn't spent a dollar on a vendor yet, but already had confirmation that users had a real problem, the data was accessible, and the scenario was technically feasible.
Insight 8: Between "Works Once" and "Works for Everyone" Is a Separate Engineering Phase
BotsCrew's job wasn't to prove the idea had value — the CEO had already done that. The task was turning the prototype into a production layer on top of NetSuite, built around how S&B Filters operates.
The rollout happened in phases. First came the Internal Order Status Assistant, which replaced manual lookups and learned to handle edge cases — including backordered items and requests where the user only provided a part number.
Next, a knowledge base was integrated, combining transactional order data with rules and product information, kept current through incremental syncing. Shopify came next, so the assistant could work not just over phone and email but in e-commerce too. After that, email and Zendesk were connected, letting correspondence link directly to order data and cutting down on manual triage.
The hardest part was cancellation and order-change logic. Here, a wrong move by the AI could carry real financial consequences and damage customer trust — so this phase required the most detailed requirements alignment in the entire project.
Each phase was stabilized before moving to the next. That's the difference between a prototype and a product: a prototype has to prove something is possible; a product has to hold up under repeated use, exceptions, multiple channels, and accountability for the outcomes.
Has to prove one thing — that it's possible.
Has to hold up under:
- Repeated use
- Exceptions and edge cases
- Multiple channels
- Accountability for the outcomes
→ Got a scrappy internal prototype that needs to become production-ready?
A weekend build from your CEO or a sharp ops lead can prove the concept — but proving it and running it for hundreds of users through phone, email, Shopify, and Zendesk are two very different jobs. BotsCrew specializes in exactly that handoff: taking a working proof-of-concept and re-engineering it for speed, stability, multi-channel support, and the kind of accountability that comes with letting AI touch real orders and real customers.
What These Projects Say About AI in Manufacturing
AI won't fix a process the company doesn't understand itself. At Kravet, the first thing to sort out wasn't the model — it was the documents: which ones were current, which duplicated each other, and how they connected to specific products. At S&B Filters, it was understanding exactly where order information lived and how it flowed through NetSuite, support, email, and Shopify.
Without that groundwork, AI just processes wrong or incomplete data faster. It can't figure out on its own which document is newer, which system takes priority, or what to do when two sources disagree.
So AI adoption doesn't start with picking a model. It starts with mapping the process itself: where requests originate, which systems hold the needed information, who's responsible for keeping it current, and at what point a decision actually gets made.
A good answer isn't automatically a good AI product. AI can answer a question correctly and still be a pain to use. If the answer takes 4–6 minutes to arrive, as with S&B Filters' first prototype, users won't experience the assistant as a faster way of working. If an employee has to bounce between several systems, copy data manually, and double-check the result by hand, AI isn't removing the busywork — it's just adding another step.
The quality of an AI product isn't just accuracy. Response speed, access to current data, ease of workflow integration, and predictability of results all matter just as much. In other words, AI shouldn't just be "smart." It has to fit naturally into how people work and make that work easier — not force the user to adapt to the technology.
AI shouldn't just be "smart." It has to fit naturally into how people work — and make that work easier.
People don't disappear — their work shifts to a different level. In manufacturing, AI can absorb the repetitive parts: looking up an order status, checking product information, pulling data from multiple systems. That doesn't mean people become unnecessary. Their role changes. A person still has to decide which sources to trust, what to do when data conflicts, and when the AI can act on its own versus when it needs to hand off to a human. People also handle exceptions and make the calls that carry financial or reputational risk.
Answering a question about order status, for instance, is a relatively low-risk scenario. Canceling or modifying that order is an entirely different level of responsibility. AI takes the mechanical work off a person's plate. Accountability for the rules, the exceptions, and the consequences still sits with people.
A successful pilot doesn't mean the system is ready to scale. The fact that an assistant handles order-status questions well doesn't mean it's ready to handle cancellations. Likewise, an internal knowledge assistant that helps employees find information isn't automatically ready to talk to customers.
Every new use case comes with its requirements for accuracy, speed, and safety. For internal search, an error might just mean an employee double-checks with a coworker or the source document. For a customer-facing assistant, the same kind of error can lead to a wrong order. And if the AI is changing or canceling orders on its own, the consequences can be financial.
Therefore, scaling isn't about copying the first successful solution across every department. Each new use case needs its own review: what data it needs, what it's allowed to do, where human confirmation is required, and what happens if the system gets it wrong.
An AI solution becomes an enterprise product not when it works in one scenario, but when the company understands the boundaries of its responsibility and can safely expand its use from there.
Questions to Ask Before Launching an AI Project
Before signing off on the next AI budget, leadership teams should be able to answer:
Will the pilot expose the weak points in our data?
A pilot can be designed in very different ways. You could take ten pre-selected questions, curate clean data for them, and show leadership the best-case answers. That kind of demo makes a good first impression but tells you almost nothing about real-world readiness.
The alternative is testing the AI on everyday work queries, contradictory documents, outdated files, inconsistent product names, and edge cases. That second approach is the one that actually informs the business. A good pilot shows you where the system gets things wrong, what data it's missing, which sources contradict each other, and which scenarios absolutely require human review.
If all you know after a pilot is that the AI "answers demo questions nicely," that's not enough. A good pilot should leave behind more than a presentation — it should leave a list of problems to solve before scaling.
Which system is the source of truth?
In a large company, the same piece of information often lives in several places at once. Product data can sit in the ERP, the PIM, the CMS, a CRM, a sales spreadsheet, and even an old PDF.
As long as those sources agree, the problem can stay invisible. But once they contain different prices, specs, statuses, or product names, AI can't just "pick the right one" without defined rules to follow.
Before launch, determine:
— which system takes priority for each type of data;
— who's responsible for keeping that information current;
— what the system should do when sources disagree;
— whether users should see which source an answer is based on.
This isn't purely a technical question. If a company hasn't defined where the current, correct information actually lives, AI can't build a reliable answer. It will just surface — faster — the fact that the organization doesn't have an agreed-upon version of the truth.
What are we going to do with the prototype that already exists?
At a lot of companies, the first AI prototype isn't built by an outside vendor — it's built by someone inside: a CEO, an IT person, an ops team, or an individual employee trying to solve their own problem quickly.
The real value of a prototype like this is that it already proves demand. The team spotted the problem, found access to the right data, and confirmed the scenario was technically doable.
But a prototype is usually built for one user, one channel, and a limited slice of data. It might be slow, unstable, dependent on manually written prompts, and lacking access controls.
Before scaling, you need to understand:
— what the prototype has proven;
— what requirements it doesn't cover;
— how many users the system needs to support;
— what happens if the data is incomplete or contradictory;
— which actions the AI can take on its own, and which need human sign-off.
A prototype is a solid starting point. But the path from "works for me" to "works for the whole team" takes a separate round of engineering.
How will we measure quality?
Showing a handful of correct answers isn't enough to draw a conclusion about an AI system's quality.
Before launch, you need to agree on what "working well" actually means. That could be answer accuracy, response time, the share of requests handled without human involvement, or how often employees still end up opening the original source anyway.
It's also important to test the system across real scenarios:
— simple and complex queries;
— different types of users;
— current and outdated data;
— standard and edge cases;
— across different time periods and after source updates.
Stability deserves its own measurement too. If the same question routes to different sources today versus tomorrow, users will stop trusting the system fast.
So evaluating AI isn't a one-time test before launch — it's an ongoing quality-control process. The model, the data, or the integrations can all change, and yesterday's result doesn't guarantee tomorrow's.
Are we accounting for every channel people actually use?
A customer might message in chat, call, send an email, or place an order through Shopify. An employee might work in NetSuite, Zendesk, a CRM, or an internal portal.
To the user, it's all the same problem. To the company, it's several different channels, integrations, and data formats.
If AI is built for just one channel, you'll eventually find it's blind to key information — or answers inconsistently — everywhere else. Then the team ends up layering on separate, disconnected fixes.
So from the start, you need to define:
— where users are actually asking questions today;
— which channels need to share the same underlying data;
— where the answer should be identical, and where it needs a different level of detail;
— where the AI should only inform, and where it should be allowed to act;
— how a request gets handed off from AI to a human when the system isn't confident.
AI is better designed around the real path of the work than around a single interface. Otherwise, you end up automating one segment of a process while leaving manual work in place everywhere else.
What Should Be Left After This Review
If there are no clear answers to these questions, that doesn't mean a company isn't ready for AI. It means they need a preparation phase — an audit of their data, processes, and integrations.
And if the answers are there, it becomes much easier to define the right scope for a pilot, its success criteria, the data sources it needs, and the boundaries of the system's responsibility.
BotsCrew has helped enterprise teams turn chaotic ERPs, scattered documents, and scrappy in-house prototypes into AI assistants their whole business can rely on.
Schedule Your Free 30-Min Consultation