Blog

AI & Machine Learning articles

Product Recommendations with Machine Learning

How a recommendation system works: simple baselines, co-purchase SQL, collaborative and content-based filtering, cold start, testing and privacy.

4 min read AI & Machine Learning

"Customers who bought this also bought", "Recommended for you", "Complete the look": product suggestions are now expected on almost any online store or content platform. Behind them sits a recommendation system, software that predicts which items a particular person is likely to want. Done well, recommendations help customers discover relevant products and can increase order value. Done badly, they show the same bestsellers to everyone or suggest a second sofa to someone who just bought one. This article explains the main approaches, from a few lines of SQL to machine learning, and how to tell whether yours is working.

Where recommendations appear

Different placements answer different questions, and often need different methods:

  • Product page: similar items, or items often bought together.
  • Basket and checkout: complementary add-ons, such as cases for a phone.
  • Home page: personalised picks based on a visitor's history.
  • Emails: replenishment reminders and items related to past purchases.
  • Search results: re-ranking matching products for the individual.

Start with simple baselines

Before any machine learning, build simple versions that are easy to understand and measure:

  • Bestsellers, overall and per category, recalculated regularly.
  • Trending: items whose sales jumped recently.
  • Merchandiser rules: hand-picked accessories for key products.

These are surprisingly effective for small catalogues or new stores, and they become the baseline any model must beat.

"Frequently bought together" in SQL

A useful first data-driven step is counting co-occurrence: how often two products appear in the same order. This query finds, for each product, the items most often bought alongside it:

SELECT a.product_id       AS product,
       b.product_id       AS also_bought,
       COUNT(*)           AS orders_together
FROM order_items a
JOIN order_items b
  ON a.order_id = b.order_id
 AND a.product_id <> b.product_id
JOIN orders o ON o.id = a.order_id
WHERE o.created_at >= CURRENT_DATE - INTERVAL 12 MONTH
  AND o.status = 'completed'
GROUP BY a.product_id, b.product_id
HAVING COUNT(*) >= 5
ORDER BY product, orders_together DESC;

(MySQL interval syntax shown.) Raw counts favour popular items: carrier bags appear "together" with everything. Normalising by popularity fixes that. A common measure is lift: how much more often two items are bought together than you would expect if purchases were independent. Run this as a scheduled job and store the top results per product in a small table the website reads quickly.

Machine learning approaches

Collaborative filtering

Collaborative filtering recommends items based on the behaviour of similar customers, without needing to know anything about the items themselves. If many people who bought A and B also bought C, someone who bought A and B may like C. Modern versions use matrix factorisation or neural models that learn a compact numerical profile for each customer and each item from purchase, click or rating data. Its strength is discovering non-obvious connections; its weakness is that it knows nothing about new items or new customers.

Content-based filtering

Content-based filtering recommends items similar in attributes to what a customer liked: category, brand, price range, colour, description. Text descriptions and images can be converted into numerical representations (embeddings) so that "similar" captures meaning rather than exact matches. It handles new products well, since attributes are known from day one, but tends to suggest more of the same.

Hybrid and ranking models

Larger systems usually combine both in two stages: first generate candidates from several sources (co-purchases, similar items, popular in category), then rank them with a model that uses customer, item and context features such as device, time and current basket. Business rules then filter the final list.

Business rules still matter

Whatever the model suggests, apply rules before showing it:

  • Do not recommend out-of-stock or discontinued items.
  • Do not recommend what the customer just bought, unless it is a consumable.
  • Respect age restrictions and regional availability.
  • Limit how many items come from one brand or category, to keep variety.
  • Allow merchandisers to pin or exclude items for campaigns.

The cold start problem

Cold start refers to new customers with no history and new products with no sales. Common answers: show popular or trending items to new visitors, use their current session (what they are viewing right now) as the signal, and use content-based similarity for new products until they gather sales data.

Measuring whether a recommendation system works

Offline tests on historical data help compare approaches. A typical method hides each customer's most recent purchases and checks whether the system would have recommended them, measured with metrics such as precision at k (how many of the top k suggestions were relevant). But offline results do not always match real behaviour.

The real test is an A/B test: show some visitors the new recommendations and others the baseline, then compare click-through, conversion and revenue per visitor over a meaningful period. Watch for side effects too, such as recommendations that increase clicks but not purchases, or that shift sales without growing them.

Pitfalls and responsibilities

  • Popularity bias: systems can keep promoting the same bestsellers, starving the rest of the catalogue of exposure.
  • Feedback loops: items get recommended because they were recommended before.
  • Privacy: recommendations reveal what is known about a person. Avoid sensitive inferences, honour consent choices for tracking, and be careful with recommendations visible to others on shared devices or in emails. Data protection rules vary by country.
  • Data quality: duplicate product records or poor category data weaken every method.
  • Cost versus benefit: for a catalogue of a few dozen products, curated suggestions may be all you need.

Our AI and machine learning development team builds recommendation engines, and our web application development team integrates them into storefronts and apps.

Key takeaways

  • Start a recommendation system with bestsellers and co-purchase SQL; they are the baseline to beat.
  • Collaborative filtering finds hidden connections; content-based filtering handles new products.
  • Apply business rules to every list, and plan for cold start.
  • Judge success with A/B tests on real purchases, not just offline metrics.

Need help with this?

Netifi helps businesses around the world with AI & Machine Learning. Tell us what you are working on.