Skip to main content
    ← Back to blog

    Machine learning: what it is and how it works

    October 04, 2026 · Equipe Sherlok
    Inteligência Artificial
    Machine learning: what it is and how it works

    Machine learning is not the hard part. The hard part is having history clean enough for a model to learn from.

    The Meta Ads auction deciding how much to pay for your next click runs on machine learning. The bank that blocked a suspicious charge on your card runs on machine learning. So does the recommendation shelf on the store you browsed yesterday. None of those were pitched internally as an "AI project". They simply had tidy enough data to train something on top of it.

    So the useful question is not whether your company should use it. It already does, in four or five places. The question is where it is running with nobody watching, and at what point it makes sense to train a model of your own.

    Machine learning is the branch of artificial intelligence in which software learns patterns from past examples instead of following rules written by a person. You show it thousands of cases whose outcome is already known, the model works out the relationship between the variables, and it then estimates the outcome of new cases.

    How a machine learning model actually learns

    The difference between traditional software and machine learning comes down to who writes the rule. In traditional software, an analyst decides: "if a customer has not logged in for 45 days, flag them as a churn risk." In machine learning, nobody picks that 45. The model sweeps the history and finds that the number separating the two groups is 38, and that it only matters when paired with a drop in usage the month before.

    The cycle has four stages.

    1. Labelled history. A table where each row is a past case and one column holds what actually happened: churned or not, bought or not, how much they spent.
    2. Training. The model tries to predict that outcome, gets it wrong, measures the size of the error and adjusts its internal weights. Thousands of times over.
    3. Validation. You hold back a slice of history the model has never seen and test on it. This is where you find out whether it learned the pattern or just memorised the table.
    4. Production. A new case arrives with no known outcome and the model returns an estimate. Weeks later the real outcome shows up and becomes training data for the next version.

    That fourth stage is the one everyone forgets. A model in production ages: customer behaviour shifts, the product mix shifts, prices shift. A model trained before a price increase keeps answering with the same confidence as always, only now it is wrong.

    How a model learns History with known outcomes Training err, measure and adjust Model validated on unseen data Prediction for a new case The real outcome that arrives later becomes training data Without that return loop, the model ages and keeps being confidently wrong.
    The four stages of machine learning, from historical table to a prediction about a new case.

    What are the types of machine learning?

    Academic taxonomies vary, but three categories cover nearly everything that shows up inside a company.

    TypeWhat it needsTypical business useWhere it usually fails
    SupervisedHistory with the outcome recorded case by caseChurn prediction, lead scoring, demand forecastingWhen the outcome was never properly filled in the CRM
    UnsupervisedJust the data, with no known outcomeCustomer segmentation, anomaly detection in ad spendIt produces clusters that are statistically valid and commercially useless
    ReinforcementAn environment where the system tries and gets feedbackBid optimisation in ad auctions, dynamic pricingIt needs a high volume of attempts before it learns anything

    In practice, almost everything an average company needs is supervised learning. And the bottleneck is rarely the algorithm: it is the outcome column. Somebody has to have recorded, case by case, what actually happened. Without that there is nothing to learn, however sophisticated the tool.

    Is machine learning the same thing as artificial intelligence?

    No. Artificial intelligence is the umbrella. Machine learning is the approach that has dominated that umbrella since the 2010s. Deep learning is a type of machine learning built on neural networks with many layers. And generative AI, including the language models that became shorthand for AI in everyday conversation, is an application of deep learning trained to produce new content.

    The distinction that matters day to day is one of purpose. A classic model answers "how likely is this customer to cancel in the next 30 days?" with a number between 0 and 1. A generative model answers "write the retention email for this customer." Those are different jobs. Conflating them explains a good share of corporate frustration with AI: people ask for a forecast from a tool built to write prose.

    Artificial intelligence any system performing tasks we associate with cognition Machine learning learns patterns from examples, with no hand-written rule Deep learning neural networks with many layers, for unstructured data Generative AI produces new text, images and code on demand
    Machine learning sits inside artificial intelligence, and generative AI is one specific application of deep learning.

    Where machine learning already runs in your business

    In the study "Unlocking AI Potential in Brazil 2026", presented by AWS in September 2026, roughly 12 million Brazilian companies say they already use artificial intelligence, yet 58% remain at the basic stage, applying it only for operational efficiency. Access stopped being the problem. Shallow use is the problem.

    Four places where a model is probably already switched on in your operation:

    • Automated bidding in paid media. Maximise conversions, target ROAS and Meta's automatic strategies are predictive models estimating the conversion probability of each impression. When you change the attribution window, you are changing the label that feeds that model. Worth reading alongside the right order for analysing Google Ads campaigns.
    • GA4 predictive metrics. Purchase probability, churn probability and predicted revenue are models trained on your own traffic. They demand volume: per Google's documentation, a property needs at least 1,000 returning users who triggered the condition and 1,000 who did not, within the last 28 days. Below that the metric simply never appears, which is why so many accounts have never seen the feature working. See also what is actually worth measuring in GA4.
    • Fraud and credit scoring. Payment gateways score every transaction in milliseconds using supervised models trained on past chargebacks. The label here is expensive and slow, because a fraud is only confirmed weeks later.
    • Demand forecasting. Retail and manufacturing use time series to estimate sales per item and per store. It works well for fast-moving goods and badly for the long tail, where history is too thin to form a pattern.

    When is it worth training your own model?

    Three questions, in this order. Does the decision repeat many times a month? Does history exist with the outcome recorded? And above all: does the prediction change a concrete action? If the answer to the third is no, stop here. A forecast that changes no behaviour is an expensive report.

    It helps to look at the order of magnitude. Picture a B2B operation with 8,000 active customers and 3% monthly churn, so 240 departures a month. Suppose the model flags 60% of them in advance: 144 customers land on the risk list each month. The sales team works that list and reverses 20% of cases, 29 customers. At an average ticket of 240 per month, that is roughly 7,000 in recurring revenue preserved monthly, near 83,000 a year. Those figures are a simulation built to show scale, not research findings.

    Notice what the arithmetic demands: somebody has to make the call. If the risk list arrives and nobody acts, the model cost money and returned a handsome report. And the whole calculation collapses when the reason for leaving is price or a cheaper competitor, because then the call reverses nothing. The model correctly names who is leaving, and the revenue leaves anyway. Anyone wanting the logic behind this will find more in predictive analytics in practice and in AI sales forecasting.

    What stalls these projects is not the algorithm

    One number shows up in nearly every machine learning presentation: "our model hit 94% accuracy." Most of the time that number means nothing. If only 3% of customers churn monthly, a model that answers "nobody will churn" for everyone is 97% right and completely useless. Accuracy on an imbalanced base misleads, and what matters is how many of the real departures the model managed to catch.

    The second problem is sneakier. It is called data leakage, and it happens when training includes information that in real life only exists after the outcome. Feeding "cancellation reason" into the variables predicting cancellation produces a flawless model in testing and a useless one in production, because at prediction time that field is still blank.

    The third is the most common and the least glamorous: duplicate records, dates in three formats, value fields mixing commas and periods. The model learns the noise along with the pattern. Worth reviewing the silent errors contaminating your analyses before training anything at all.

    Frequently asked questions

    Is machine learning the same as artificial intelligence?

    No. Artificial intelligence is the whole field, and it also includes rule-based systems. Machine learning is the set of techniques where a system learns from data, and it accounts for most of what the market calls AI today.

    What is the difference between machine learning and deep learning?

    Deep learning is a subset of machine learning built on neural networks with many layers. It shines on unstructured data such as images, audio and text. For tabular data, which is what most companies have, simpler models often deliver equal or better results at lower cost and with far more explainability.

    Do I need to code to use machine learning?

    To train a model from scratch with fine-grained control, yes, usually in Python. To use machine learning in business decisions, no. Ad platforms, CRMs and analytics tools already embed trained models, and it is now possible to ask questions about your own database in plain language and get the analysis back.

    How much data is needed to train a model?

    It depends on the problem, but the useful frame is counting examples of the rare event, not total rows. A million orders with only 40 frauds is a poor base for fraud detection. As a concrete benchmark, GA4 only switches on its predictive metrics with 1,000 users on each side of the condition within 28 days.

    Does machine learning work for a small company?

    It works as long as the volume of decisions justifies it. A company with 200 customers does not need a churn model: the owner knows all 200 by name. The tipping point arrives when nobody can review cases one by one any more and decisions start being guesses.

    Where to start

    The path that tends to work is the opposite of what people imagine. Instead of picking an algorithm and hunting for somewhere to apply it, pick a decision you make blind every week and check whether its history was ever recorded. In most companies the answer is uncomfortable: the history exists, but it is scattered across the ad platform, the CRM and three spreadsheets that never talk to each other.

    Fixing that part pays better than any sophisticated model. With the data gathered in one place, a good share of the questions that would justify a machine learning project answer themselves, by asking in plain language and getting the analysis back, which is exactly what Sherlok does. The predictive model comes later, once the question is sharp and the decision has an owner.

    Want to see this in your own data?

    Connect your accounts and ask your first question in 5 minutes.

    • No card to start
    • Nothing changes in your accounts without your approval

    Keep going

    Read next