Retention

Next best action

Next best action is a decision method that picks, for one customer at one moment, the single contact with the highest expected value from a set of possible actions, after rules and contact limits have removed the ones that are not allowed.

In short

Next best action (NBA) is a decision method that chooses, for each customer at each moment, the one contact most worth making. It filters out actions the customer is not eligible for, scores the rest by likelihood and value, and applies contact limits before anything is sent. The best systems score the change an action causes, not just the chance of a response.

Origin
CRM and decisioning practice; no single inventor (Pega and Salesforce document the common form), no agreed date
Level
401 · Expert
Fits
Enterprise
Time to apply
2 to 4 weeks for a rules-and-value first version on one channel
What you need
a list of the actions you might take, each with an owner · customer data that can answer eligibility questions · a response log that records what each customer was shown and did next · a holdout group that receives no contact, even for a small test

Next best action (NBA) is a decision method that picks, for one customer at one moment, the single contact most worth making. It grew out of CRM and machine-learning practice rather than from one author, so there is no agreed inventor or date. What exists is vendor documentation that describes the same shape: Pega’s Customer Decision Hub and Salesforce’s Einstein Next Best Action both turn a long list of possible actions into one or a few suggestions. The sections below follow the decision in order, then cover the two bodies of research that improve it.

How does a next best action decision work?

A decision has four stages: rules remove actions the customer may not receive, a model estimates how likely each remaining action is to work, a value turns that into a ranking, and contact limits decide whether the winner may go out. The picture shows the order.

Five boxes joined by arrows from left to right: Candidate actions, Eligibility rules, Score and rank in blue, Contact limits, One action sent.
Rules remove options, scoring ranks what is left, and contact limits decide whether the winner may go out.

The placement of contact limits varies by product, so treat the order as the common logic, not a standard. The sections that follow take each stage in turn.

Eligibility: what may this customer receive?

Eligibility is the set of hard rules that decide which actions a customer can be offered at all. Pega’s training material calls them strict rules for what is legal, and even possible, to offer, such as a minimum age for a credit card.

Pega splits its engagement policies into three layers and applies them in a fixed order:

Layer Question Example from Pega
Eligibility Is it allowed and possible? Gold Card requires age 18 or over
Applicability Does it make sense given what they hold? Customer already has a Platinum Card
Suitability Is it appropriate for this person? Debt-to-income ratio below a threshold

The split matters because the layers fail differently. An eligibility mistake is a compliance problem. A suitability mistake is an annoyed customer. Keep all three out of the scoring step, so no large number can outvote a rule.

Propensity, value and arbitration

Propensity is the predicted likelihood that a customer responds positively to an action, a number between 0 and 1. In Pega it comes from adaptive models that update as outcomes arrive, using either a Bayesian method or gradient boosting.

Value is what a successful outcome is worth in money, net of cost. Probability alone misleads: a tip with a 30% chance and a value of 5 loses to an upgrade with a 4% chance and a value of 400.

Arbitration is the step that combines these into one ranking. Pega’s formula multiplies propensity (P), context weighting (C), business value (V) and business levers (L), and picks the highest product. Context weighting lifts an action that fits the moment, such as retention when a customer calls to close an account. Levers let the business push toward a goal. Salesforce takes a different route: its strategies filter and sort recommendation records with Flow Builder elements, so the ranking rule is yours to write.

Contact policy: when to stay quiet

A contact policy limits how often a customer hears from you. Pega describes two types. Contact limits cap how many actions a customer receives on a channel in a period, for example two promotional emails a month. Suppression policies pause an action after repeated outcomes: ignore a card promotion twice in 30 days and it is held for the next 60.

These limits are the reason NBA can answer “do nothing”. A trigger flow fires whenever its event happens. An NBA engine sees every action competing for the same customer and can let all of them wait.

Propensity is not enough: the uplift problem

Propensity measures who will act. It does not measure who acts because of you. Nicholas Radcliffe and Patrick Surry argued in 2011 that most targeted marketing is aimed with non-incremental models, even when results are judged by incrementality. Uplift modelling estimates the change a contact causes, and it needs a control group to do so.

A two-by-two grid. The horizontal axis is Buys if not contacted and the vertical axis is Buys if contacted. Top left is Persuadables in blue, top right Sure things, bottom left Lost causes, bottom right Sleeping dogs.
Only persuadables change their behaviour because of the contact.

The grid shows why. Sure things buy anyway, lost causes never buy, and sleeping dogs buy less if contacted. Radcliffe and Surry note that retention activity can backfire for some customers, and a propensity model cannot see it. The same paper gives a sizing rule of thumb: the product of the overall uplift and the size of each group should be at least 500, so a 0.1% uplift needs 500,000 customers in both treated and control groups.

Marketing research points the same way. Eva Ascarza found in field data from a wireless provider and a membership organization that targeting by lift cut churn more than targeting by risk of leaving. Lemmens and Gupta added customer profit to the loss function and showed in two field experiments more profitable campaigns than competing models. For the methods, Gutierrez and Gérardy review the uplift literature, and Hitsch, Misra and Zhang found that direct estimates of treatment effects beat indirect ones in a catalog-mailing test. Measuring the effect in the first place is the subject of incrementality testing.

Contextual bandits: learning while deciding

A contextual bandit is an algorithm that picks one action given what it knows about the customer, sees the result and updates, trading off trying new actions against repeating the best one. It suits NBA because the system must keep choosing while it learns.

Li, Chu, Langford and Schapire used it for news articles on Yahoo’s front page and reported a 12.5% click lift over a context-free bandit on more than 33 million events. Schwartz, Bradlow and Fader ran a bandit in a display campaign for a bank across 750 million impressions and raised acquisition by 8% over a control. They also estimated that optimising clicks instead of conversions would cut acquisition by about 10%, a warning about choosing the outcome. Before going live, a team can test a policy on logged data with replay evaluation or doubly robust estimates.

How NBA differs from its neighbours

Method Decides Needs
RFM analysis Which customers are most valuable Purchase history
Trigger emails What to send after an event An event and a message
Customer lifecycle marketing What each stage needs A stage model
Next best action Which one action, now, among many Rules, scores, values, limits

NBA does not replace the others. Lifecycle stages and RFM segments make useful inputs to the scoring, and triggers are often among the candidate actions. A Growth Lab plan starts from the actions and rules, because the model is the easy part.

How to apply Next best action, step by step

  1. List the actions and their rules. Write every action the business could take with a customer: a renewal call, a discount, a service tip, doing nothing. Give each a hard rule for who may receive it, such as age, consent or product held. Result: a catalogue where every action carries its own eligibility test.
  2. Choose one outcome and estimate propensity. Pick the outcome that matters for each action, such as accepting an offer within seven days. Estimate for each customer how likely it is. A simple model on past responses is enough to start. Result: a probability between 0 and 1 for every customer and action pair.
  3. Attach a value to each action. Put a money value on a successful outcome, net of what the action costs. A renewal worth 400 and a tip worth 5 should not be ranked on probability alone. Result: one expected-value number per action.
  4. Test the incremental effect. Hold back a random share of customers from each action and compare outcomes. Replace raw propensity with the measured difference where you can. Result: evidence that an action changes behaviour, and a list of customers for whom it does not.
  5. Set the arbitration rule and the contact limits. Decide how scores combine into one ranking, then cap how often any customer is contacted and how long an ignored action is paused. Result: a written rule that returns one action, or none, for any customer on any day.
  6. Run it, log it and keep a holdout. Send the winner, record what happened, and refit the scores on the new data. Keep a permanent holdout so the whole system is measured against doing nothing. Result: a decision engine that can show its own value.

Examples

A dental clinic with one call to make

Illustrative. The front desk has time for 40 calls a day. Candidate actions are a recall for a patient overdue for a check-up, a follow-up after a filling and a payment reminder. Eligibility removes patients who declined calls. Scoring ranks the rest by the chance the call books a visit, multiplied by the value of that visit. A contact rule stops a patient being called twice in a week. The desk sees a ranked list instead of a spreadsheet.

A fintech app choosing among three messages

Illustrative arithmetic. Offer A, a savings tip: 10% chance of acceptance, worth 100 if accepted, so expected value 10. Offer B, a premium card upgrade: 4% chance, worth 400, so expected value 16. Offer C, a card-limit reminder: 30% chance, worth 5, so expected value 1.5. Ranking by probability picks C. Ranking by expected value picks B, unless a rule such as a credit check makes the customer ineligible, in which case A wins.

When to use it

Use it when you have several possible contacts, a good response log and enough customers that one person cannot choose by hand. It fits banks, telecoms, insurers, clinics with recall programmes and subscription products where the same customer can receive many different messages.

When not to use it

Skip it when you have one product, one channel or a few hundred customers: a simple trigger flow or a segment list does the job. Skip it where regulation fixes who must be contacted and when. Without a response log and a holdout, the scoring step has nothing to learn from.

Common mistakes

  • Ranking by propensity alone. A customer who would have bought anyway scores high and gets a discount for nothing. Rank on incremental effect times value where you can measure it.
  • Treating a rule as a score. Eligibility is a yes or no and belongs before the ranking. Mixing legal or consent rules into a weighted score lets a big number override them.
  • No contact limits. Each action looks good alone, and together they flood the customer. Cap contacts per channel and pause ignored actions.
  • Never exploring. A model that always sends the current best action never learns whether a new one is better. Reserve a small share of contacts for testing.
  • Skipping the holdout. Without a group that receives nothing, you cannot tell whether the engine adds anything over the old calendar.

FAQ

What is next best action in marketing?

Next best action is a method that decides which single contact to make with a customer right now. It checks which actions the customer may receive, scores them by likelihood and value, applies contact limits and returns one winner. It can recommend an offer, a service message, a call, or no contact at all.

What is the next best action formula?

There is no universal one. Pega's documentation multiplies four numbers: propensity, context weighting, business value and business levers, and picks the highest product. Salesforce does not prescribe a formula; its strategies filter and sort recommendation records with rules, predictive models and other data, and you define the order.

What does a next best action tool tell a manager or agent?

It shows a short list of recommended actions at the moment of an interaction. In Salesforce each recommendation has a description, button text and a linked flow, and the user can accept or reject it. The response is logged and can feed back into the scoring, so the suggestions improve.

What is the difference between propensity and uplift?

Propensity is the chance a customer does something, such as buying. Uplift is the change in that chance caused by contacting them. A customer with a 60% propensity who buys anyway has near-zero uplift. Uplift needs a control group, because the customer who was not contacted is never observed after contact.

How do contextual bandits relate to next best action?

A contextual bandit is an algorithm that picks one action given customer context, observes the result and updates, balancing trying new actions against repeating the best-known one. It is a way to run the scoring step while it keeps learning, and research on news and display advertising shows it can beat fixed rules.

Sources

  1. Pega Academy, Action arbitration
  2. Pega Academy, Customer engagement policies
  3. Pega Academy, Contact policy types
  4. Pega Academy, Adaptive models
  5. Salesforce Help, Suggest Options with Recommendation Strategies
  6. Salesforce Help, Get Started with Einstein Next Best Action
  7. Nicholas J. Radcliffe, Patrick D. Surry, Real-World Uplift Modelling with Significance-Based Uplift Trees, Stochastic Solutions, 2011
  8. Pierre Gutierrez, Jean-Yves Gérardy, Causal Inference and Uplift Modelling: A Review of the Literature, PMLR 67, 2017
  9. Piotr Rzepakowski, Szymon Jaroszewicz, Decision trees for uplift modeling with single and multiple treatments, Knowledge and Information Systems, 2012
  10. Susan Athey, Guido Imbens, Recursive Partitioning for Heterogeneous Causal Effects, arXiv 1504.01132
  11. Sören Künzel, Jasjeet Sekhon, Peter Bickel, Bin Yu, Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning, arXiv 1706.03461
  12. Eva Ascarza, Retention Futility: Targeting High-Risk Customers Might Be Ineffective, Journal of Marketing Research, 2018 (American Marketing Association award notice)
  13. Aurélie Lemmens, Sunil Gupta, Managing Churn to Maximize Profits, Marketing Science 39(5), 2020
  14. Günter Hitsch, Sanjog Misra, Walter Zhang, Heterogeneous treatment effects and optimal targeting policy evaluation, Quantitative Marketing and Economics 22(2), 2024
  15. Lihong Li, Wei Chu, John Langford, Robert Schapire, A Contextual-Bandit Approach to Personalized News Article Recommendation, WWW 2010
  16. Lihong Li, Wei Chu, John Langford, Xuanhui Wang, Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms, WSDM 2011
  17. Olivier Chapelle, Lihong Li, An Empirical Evaluation of Thompson Sampling, NIPS 2011
  18. Daniel Russo and others, A Tutorial on Thompson Sampling, Foundations and Trends in Machine Learning, 2018
  19. Shipra Agrawal, Navin Goyal, Thompson Sampling for Contextual Bandits with Linear Payoffs, arXiv 1209.3352
  20. Aleksandrs Slivkins, Introduction to Multi-Armed Bandits, Foundations and Trends in Machine Learning, 2019
  21. Miroslav Dudík, John Langford, Lihong Li, Doubly Robust Policy Evaluation and Learning, ICML 2011
  22. Eric Schwartz, Eric Bradlow, Peter Fader, Customer Acquisition via Display Advertising Using Multi-Armed Bandit Experiments, Marketing Science 36(4), 2017

Last updated Oct 9, 2026

Ilia PushinFounder, PUSHERS & COO Fintech ServiceIlia builds operating systems for growing companies in fintech and healthcare. Since 2021 he has run cross-border payments at ARBI Exchange, a licensed currency exchange in Thailand, including KYC and AML and the move into new jurisdictions.About the authorLinkedIn
Related frameworks
More frameworks
Want Next best action running inside your company?Request an operations audit