Vendors tend to describe AI as if it returns roughly the same value everywhere. It does not. The determining factors are how much repeated decision volume a sector has, how much of its data is already captured, and how expensive its errors are. Those vary enormously, and they determine whether a deployment pays back in four months or never.
Here is our honest read on the seven sectors we work in, including where we think the returns are weaker than the marketing suggests.
RMG & apparel: strong, and under-exploited
Garment manufacturing has close to ideal conditions: enormous repetition, tight margins, and errors that are expensive precisely because they are caught late. Vision-based defect detection at the sewing station rather than at final QC recovers labour that is otherwise simply lost, and the volume means even a modest per-unit saving compounds hard.
The constraint is rarely technical. It is that floor data is often not captured at all. Production is tracked on paper or in a spreadsheet reconstructed after the shift. That makes the first project as much an instrumentation project as a modelling one, which is worth knowing before you budget it.
Healthcare: strong, but not where people expect
The public conversation is about diagnosis. The reliable returns are in documentation, scheduling, coding and triage, the administrative mass that pulls expensive clinicians away from patients. Ambient note-drafting reliably gives clinicians a meaningful part of their day back, and no-show prediction recovers capacity that was already fully paid for.
Diagnostic AI is genuinely valuable but carries a regulatory and liability burden that changes the project shape entirely. We generally recommend earning the organisational trust on administrative wins first.
Security & surveillance: strong, with a caveat
Most sites already have complete camera coverage and no realistic capacity to watch it. Turning passive recording into active detection is one of the clearest value propositions available, because the capital expenditure has already happened and the incremental cost is inference.
The caveat is that false-positive tolerance is far lower than teams expect. An alerting system that cries wolf twice a shift gets muted within a week and is then worse than nothing, because everyone believes they are covered. Tuning the threshold is the actual project.
Retail & e-commerce: strong on demand, mixed on support
Demand forecasting and markdown optimisation are dependable earners. The data exists in your order tables, the decisions repeat constantly, and the savings land directly in working capital and realised margin.
Support automation is more variable than the category's reputation suggests. Agents grounded in real order and shipping data genuinely resolve routine tickets. Agents grounded in nothing produce deflection metrics and customer frustration, which is why the honest version of this project starts with the integrations rather than the chat interface.
Logistics: strong, and the easiest to prove
Routing and load planning are classical optimisation problems with a measurable objective and immediate feedback. Of everything we deploy, this category has the cleanest before-and-after: kilometres driven and fill rate are unambiguous, so the value is not a matter of interpretation.
The realistic obstacle is driver adoption. A route a driver does not trust is a route they will not follow, and an optimiser that ignores local knowledge gets overridden into uselessness.
Finance & banking: strong, constrained by explainability
Fraud detection, alert triage and document-heavy onboarding all have excellent economics. The volumes are enormous, the data is clean by the standards of other sectors, and the cost of each error is well understood.
The binding constraint is not accuracy but explainability. A model that improves outcomes but cannot produce reason codes for a declined applicant is unusable regardless of its performance, so the architecture has to be chosen with model risk review in mind from the first week rather than retrofitted.
Manufacturing: strong on downtime, slower on the rest
Predictive maintenance is the flagship case for good reason: unplanned downtime is expensive, sensor data is often already being collected and discarded, and converting an emergency into scheduled work has an obvious value.
Visual quality inspection is equally sound where units are high-volume and defects are visually distinctive. It gets harder, and often uneconomic, for low-volume, high-mix production where you cannot gather enough examples of each defect type. That is a case where our report would tell you not to build it.
The pattern underneath
Across all seven, the sectors where AI pays best share the same profile: high decision volume, data already being generated as a byproduct of operations, and errors that are expensive because they are discovered late. Where any of those three are missing, returns drop sharply, and no amount of model sophistication compensates.
The technology is rarely the variable. Where you point it almost always is.
Which is why the useful question is never 'should we use AI'. It is 'which of our processes has volume, data and expensive lateness at the same time', and that question can only be answered by looking at your operation, not at a sector average.