Industry AI

AI for Retail and E-commerce Operations

Workisy Team
July 24, 2026
9 min
AI for Retail and E-commerce Operations

Retail is a business of small percentages compounding across enormous transaction volumes. A one-point improvement in forecast accuracy, a markdown taken two weeks earlier, a store scheduled to actual traffic rather than to last year's traffic — none of these are dramatic on their own. Together they are the difference between a category that clears at full price and one that funds the clearance rack.

What makes retail unusually well-suited to AI is not that the sector is data-rich, though it is. It is that retail decisions are repetitive, high-frequency, and consequential in aggregate rather than individually. Nobody agonizes over how many units of a mid-tier SKU to allocate to store 214. That decision gets made ten thousand times a week by a planner with a spreadsheet and an hour, and the cumulative error shows up four months later as inventory sitting in the wrong buildings.

This is a look at where AI actually changes retail and e-commerce operating numbers — organized around the path a unit takes from forecast to allocation to sell-through to return — plus the parts of the business where the technology is currently oversold.

The Forecast Problem Is Rarely the Algorithm

Most retailers who describe a forecasting problem do not have a modeling problem. They have a hierarchy and calendar problem.

Forecasts are typically produced at a level of aggregation that is convenient for planners — category by region by month — and then spread down to SKU by store by week using ratios that were set years ago and never revisited. The model may be excellent at the top. The spread destroys the accuracy on the way down, and the store-level replenishment system, which is the only place the forecast actually touches inventory, inherits the damage.

The second failure is calendar handling. Retail demand is driven by events that shift: a holiday that moves week-over-week between years, a promotion that ran in week 14 last year and week 11 this year, a competitor's clearance event, a school term that started early in one region. A model trained on raw weekly history without event alignment will confidently reproduce last year's noise.

Machine learning forecasting earns its keep by generating at the level where decisions are made, using demand rather than sales as the target — sales are censored by stockouts, and a model trained on them learns to under-forecast exactly the items that sold out — and by treating promotions, weather, price, and local events as explicit inputs rather than after-the-fact adjustments. The lift over a well-tuned statistical baseline is usually modest at the aggregate level and substantial at the SKU-store level, which is precisely where the money is.

Allocation, Replenishment, and the Cost of Being Roughly Right Everywhere

Once you have a store-level demand signal, the operational question becomes where to put the units. Retailers habitually spread inventory evenly and then spend the season transferring it, absorbing the cost of the transfer twice.

The higher-value application is size and attribute-level allocation. A model that learns store 214 skews two sizes larger than the chain average, or that a particular color performs in coastal stores and dies inland, prevents a specific and expensive failure: a store that is technically in stock on a style but out of stock in every size a customer would buy. Aggregate inventory looks healthy; sell-through does not.

Replenishment benefits from a different capability — recognizing phantom inventory. Perpetual inventory records drift from reality through theft, misscans, and receiving errors, and the symptom is a SKU that shows three units on hand and has not sold in five weeks. A model that flags these patterns for cycle count triggers more found revenue than most demand-side improvements, because the item was never actually available to sell. Retailers running integrated stock and billing systems have an advantage here, since the reconciliation between recorded and observed movement is already instrumented; the fundamentals are covered in this guide to inventory and billing management.

Markdown Timing Beats Markdown Depth

Most retailers markdown too late and too deep. The instinct is to protect margin by holding price, then to panic when the season closes and cut 50% on inventory that would have cleared at 20% six weeks earlier.

Markdown optimization models the trade-off explicitly: expected sell-through at each price point, remaining weeks in the season, salvage value at season end, and the cost of the space the units occupy. The output is a schedule rather than a single decision — take 15% in week six, 30% in week ten, exit by week fourteen — and it is revised weekly as actuals come in.

The organizational obstacle is more common than the technical one. Markdown authority usually sits with merchants who are measured on gross margin percentage, a metric that a well-timed early markdown will make look worse in the short term even as it improves gross margin dollars. Any markdown model deployed without changing that incentive will be overridden until it is quietly switched off.

Scheduling Stores Against Traffic, Not Against History

Store labor is typically the largest controllable cost after inventory, and it is scheduled with the least data. The common method is a sales-per-labor-hour target applied to a forecast that arrives as a weekly number, which the manager then splits into shifts by intuition.

The improvement is forecasting traffic and transaction volume in fifteen or thirty minute increments and translating that into a labor requirement curve that reflects what staff actually do — a fitting room needs coverage proportional to try-ons, a register needs coverage proportional to transactions, and receiving needs coverage proportional to inbound cartons, which correlates with nothing on the sales side. Scheduling to a single blended curve overstaffs the floor in the morning and understaffs it at the 5 p.m. peak.

Constraint handling is where these systems succeed or fail. A schedule that is theoretically optimal but violates minimum shift lengths, predictive scheduling notice requirements, availability commitments, or minor-employment rules is worse than a mediocre one, because the manager will rebuild it by hand and stop trusting the system. The operational groundwork for this is covered in more depth in this overview of workforce management software, and the same demand-driven staffing logic underpins the retail and hospitality solutions most multi-site operators start with.

Returns: The Line Item Nobody Forecasts

E-commerce return rates in apparel routinely run several times higher than store rates, and the cost is not the refund. It is the reverse logistics, the inspection labor, the repackaging, and the fact that a returned unit typically re-enters inventory too late in the season to sell at full price.

AI touches returns in three places. Prediction at the point of order — flagging carts with a high probability of return, often driven by bracketing behavior where a customer orders three sizes intending to keep one — allows intervention through better fit guidance rather than through blocking the sale. Automated disposition routing at the return center decides whether a unit goes back to stock, to outlet, to a liquidator, or to disposal, a judgment currently made by an inspector under time pressure with no visibility into downstream demand. And return-reason classification from free-text customer comments feeds product and vendor quality signals that would otherwise never reach a buyer.

The third is the most undervalued. A style with a 40% return rate concentrated on "smaller than expected" is a specification defect, and it is fixable before the next buy.

Personalization Past the Recommendation Widget

Product recommendations are the most visible AI in retail and among the least differentiating; the delta between a competent recommender and an excellent one is real but bounded. The larger opportunities sit in less visible places.

Search relevance is one. A meaningful share of e-commerce sessions include a search, and searches that return nothing or return irrelevant results convert at close to zero. Handling synonyms, misspellings, and attribute queries — "waterproof jacket under 100" — is a comprehension problem that modern models handle far better than keyword matching.

Promotion targeting is another. Blanket discounting subsidizes customers who would have paid full price. Modeling incremental response, meaning the lift a specific offer produces for a specific segment rather than the raw redemption rate, routinely reduces promotional spend while holding revenue flat.

Peak Changes Every Assumption

Models trained on fifty weeks of ordinary trading behave badly in the two weeks that generate a disproportionate share of annual profit. Peak has different basket composition, different fulfillment constraints, different labor mix with temporary staff, and different failure modes.

Practical guardrails matter more than model sophistication during peak. Cap the magnitude of automated changes, freeze model retraining in the weeks leading up to the event, keep a manual override path that does not require an engineer, and instrument the system so that a planner can see what changed and why. The worst peak outcomes come from automation acting confidently on a pattern it has never seen.

A Sensible Order of Operations

Start where the data already exists and the decision already repeats. In most retailers that is demand forecasting at store-SKU level or store labor scheduling, both of which have clean baselines and unambiguous measurement. Markdown optimization follows once forecast accuracy is trusted, because markdown decisions inherit forecast error. Returns disposition and personalization come later; they depend on cleaner product and customer data than most retailers have on day one.

Measure against the decision, not the model. Forecast accuracy in isolation is a vanity metric. The numbers that matter are in-stock rate on top sellers, sell-through at full price, markdown as a percentage of sales, sales per labor hour, and units touched twice in the return center.

Retailers who treat this as connected workflow automation rather than a set of isolated point tools tend to see returns sooner, because the forecast, the allocation, the schedule, and the markdown are the same decision viewed at four different moments. If you want an outside read on which of those four to instrument first, a short operational review is usually enough to identify it.

Share:LinkedInX

See These Insights in Action

Discover how Workisy can help you implement these strategies and transform your HR operations.

Request a Demo