# LoyalFlow > Practical insights on customer loyalty programs, retention strategy, and lifecycle marketing. --- # Customer Beta Testing: Pay for Research, Not Loyalty https://loyalflow.cc/blog/customer-beta-testing-research-protocol The short version: Customer beta testing buys structured product evidence, not loyalty. Pay for defined tasks, size attempts against the defects you need to detect, keep incentives outside loyalty reporting, then treat repeat use as a separate measurement problem. Key takeaways Pay a fixed $5–$15 for a short test; never vary compensation by sentiment. Allocate attempts by risky journey, not every possible demographic cross-section. Use defect prevalence to set sample size: a 5% defect needs 59 relevant attempts for a 95% chance of observation. Exclude incentivized test orders from reuse metrics; mature both comparison cohorts for 30 days. Reconcile incentives, payments, charges, and orders with deterministic rules. Customer beta testing needs a research contract A beta participant can praise an app, accept $10, then never order again. The payment bought attention and task completion. It did not buy retention, establish preference, or prove that the product change caused later behavior. The payment closes the task, not the customer relationship. Define the contract before recruitment: eligible customers, tested platform, required journeys, evidence requested, payment, start date, end date, and decision owner. A 7–14-day window works for a short consumer-app beta because it gives participants several opportunities without turning the exercise into an open-ended panel. Pay a fixed $5–$15 for up to 30 minutes . Raise that amount when testing requires purchases, multiple sessions, specialist users, or screen recordings. Reimburse required spending separately; otherwise a nominal $10 payment can become negative compensation after delivery fees or travel. Prompts should avoid suggesting the desired answer. Ask “What happened after you selected checkout?” rather than “How easy was checkout?” The first wording reduces pressure to agree with the researcher while preserving room for positive, negative, or mixed evidence. The classic failure: awarding points, prize entries, or extra payment for a five-star review. That changes the task from finding defects to producing approval. Fix compensation before the test and state that criticism cannot reduce payment. Size the beta around defects, not segments Do not divide 50 testers across platform, lifecycle stage, fulfillment method, payment type, market, and order frequency. Those dimensions create more crossed cells than the sample can support. A cell containing 8–10 people may expose an obvious failure, but it cannot reliably detect a rare one. Rare defects require wider coverage, not thinner slices. Set a target defect rate and observation probability for each critical journey. The probability of seeing at least one defect across n independent relevant attempts is 1-(1-p)^n , where p is the assumed defect rate. For a 95% observation chance, a 5% defect requires 59 attempts; a 1% defect requires 299. These are attempts, not recruited customers. One tester can provide several attempts only when the attempts are genuinely separate opportunities for the defect to occur. Repeating the same failed checkout on one device does not provide independent coverage of markets, payment processors, or operating systems. Allocate coverage by journey first: authentication, store selection, basket restoration, offer application, fulfillment changes, payment, cancellation, and help. Add explicit platform or customer strata only where the underlying state differs. Saved-card customers deserve separate coverage when token migration changed; arbitrary age bands do not unless the test has a reason to expect different behavior. A beta still cannot certify the absence of defects. If zero failures appear in 59 attempts, that does not prove a zero failure rate. Report the attempts, observed failures, journey, platform, and exposure conditions so the release owner can judge residual risk. The classic failure: declaring “no payment issues” after ten successful attempts. If the true defect rate were 1%, ten independent attempts would have only about a 9.6% chance of observing at least one failure. The test barely challenged the claim. Measure orders and reuse with valid denominators Instrument each funnel before testing. Checkout completion uses started checkouts as its denominator; payment failure uses payment attempts; support-contact rate uses eligible attempts or completed orders, stated explicitly. Counts without exposure cannot distinguish a widespread defect from a heavily used feature. Keep the paid trial in the record—and out of the denominator. Track checkout completion, technical failures, median ordering time, support contacts per 100 attempts, duplicate charges, and orders missing after successful payment. Detect duplicate charges by reconciling payment-provider transaction IDs against order IDs and charge states. Beta comments can flag the symptom; they cannot perform the reconciliation. Predeclare release thresholds. One workable structure is checkout completion no more than 2 percentage points below the current flow, zero unreconciled duplicate charges, zero inaccessible controls blocking purchase, and median ordering time within 10% of baseline. These are operating tolerances, not universal benchmarks; tighten or loosen them using order value, traffic, customer harm, and rollback cost. Thirty-day reuse requires an eligible denominator and a mature observation window. Exclude incentivized test orders, define eligibility on the same date, and wait until every included customer has had 30 full days to return. Report the absolute reuse-rate difference with a 95% confidence interval rather than presenting the point estimate alone. For a causal relaunch claim, use randomized staged access where operationally possible. Assign eligible customers within the same platform, market, prior-order band, and promotion rules to old or new experiences; predeclare a non-inferiority or lift threshold; analyze assignment rather than voluntary adoption. If randomization is unavailable, match on prior frequency, recency, market, platform, and promotion exposure, then label the result observational because unmeasured selection remains. The classic failure: comparing enthusiastic beta volunteers with all prelaunch customers. Different prior frequency, promotion exposure, store availability, and self-selection can impersonate product improvement. A baseline supplies context; it does not isolate causality. Keep incentives outside loyalty economics Book beta compensation to research or product, not loyalty rewards expense. Maintain a separate ledger containing research_id , offer date, completion status, amount, payment date, reversals, and expiry. Reconcile issued, paid, expired, reversed, and outstanding amounts arithmetically. Reward the research; leave loyalty accounting untouched. A fixed-value payment avoids point valuation, earn-rate, redemption, and breakage attribution. If points are operationally unavoidable, apply a distinct reason code such as beta_research , publish any 30–90-day expiry before participation, and exclude those points from campaign ROI and organic earning reports. Use one incentive per verified participant, then check duplicate account, payment instrument, phone, address, and device signals where lawful. Route ambiguous household matches to review rather than blocking automatically. The minimum control set is covered in Loyalty Program Fraud Prevention: Six Minimum Controls . If a vendor receives customer IDs, contact details, order history, recordings, or device data, tell participants which fields leave your systems and who receives them. Send only what recruitment, payment, and analysis require; substitute an internal research ID when direct identity is unnecessary. The classic failure: issuing ordinary bonus points without a separate reason code. Research spending then inflates issued currency, changes redemption timing, and appears as loyalty activity despite measuring product usability. Convert confirmed failures into testable acceptance criteria, owners, and release gates. Loyalty Program Software: Buy for Requirements, Not Features provides the adjacent procurement discipline. Source: www.marketingdive.com Frequently asked questions How many beta testers are enough? No universal count works. Choose the defect rate worth detecting, required observation probability, and number of independent relevant attempts. Use 1-(1-p)^n ; recruit enough customers to produce those attempts across the states that can change the result. Can beta participation measure loyalty? No. Participation measures willingness to test under the offered terms. Measure later ordering separately, exclude incentivized orders, mature the observation window, and use randomized staged access when making causal claims. Should loyalty members join the beta? Yes, when they represent the tested journey. Stratify by prior frequency because experienced members know the existing workflow; compare their task results with newer or lower-frequency customers rather than assuming either group is representative. --- # Customer Loyalty Program: Define the Repeat Behavior First https://loyalflow.cc/blog/customer-loyalty-program-repeat-behavior The short version: A customer loyalty program needs a precise behavior contract before it needs rewards or software. Define the eligible customer, qualifying event, start timestamp, observation window, identity rules, and reversals; otherwise repeat-rate reporting will not survive scrutiny. Key takeaways Specify one repeat event using fields that data systems can evaluate deterministically. Anchor the observation window to an eligibility timestamp, not enrollment or campaign delivery. Keep every eligible customer in the denominator, including non-buyers and unredeemed offers. Resolve identity, cancellations, refunds, and role-versus-purchase cases before launch. Estimate lift only after checking sample size, assignment integrity, and complete observation time. Write the customer loyalty program behavior contract “Increase engagement” is not a program objective. Neither is “drive loyalty.” A usable objective names one population, one observable event, and one deadline: first-time lunch buyers complete a second paid lunch order within 30 days of their first completed order. Loyalty starts when one exact action rings the bell. Turn that sentence into a data contract. At minimum, define customer_id , eligible_at , event_at , event_type , order_status , and net_revenue . Specify exact accepted values: for example, only order_status=completed qualifies; pending, cancelled, fully refunded, test, and staff orders do not. The qualifying event must represent the behavior being funded. If the objective is another paid order, an app open, coupon save, email click, or points balance increase cannot substitute for it. Those events may diagnose the path, but they do not satisfy the contract. The classic failure: the team defines “repeat customer” after seeing the report. One analyst counts any second order; another excludes refunds; a third starts the clock at enrollment. Freeze the definition before exposure, version any later change, then rerun treatment and comparison groups under the same version. Anchor the window and denominator correctly The window starts when the customer becomes eligible for the target behavior. For a second-purchase program, that may be the first order’s completion timestamp. Campaign delivery is not the correct anchor when messages arrive hours or days later, because it gives customers different behavioral windows. Count the quiet customers too. Define boundary handling explicitly. A 30-day window can use event_at>eligible_at and event_at<=eligible_at+30d . Pick one timezone, document whether the final boundary is inclusive, then apply the same rule in campaign selection and reporting. The repeat-rate denominator is every eligible customer whose full outcome window has matured. If 800 first-time buyers became eligible and 144 completed the event, repeat rate is 18%. Do not divide by members who opened the email, activated the offer, or returned to the site; those filters remove non-responders and inflate the result. Late entrants need time to mature. Someone eligible yesterday cannot yet be classified as a 30-day non-repeater. Either wait until the full cohort completes its window or report fixed entry cohorts separately, with matured_eligible as the denominator. Denominator trap: a dashboard silently excludes customers with no later event because its query begins from the purchase table. Build the eligible population first, then left-join qualifying events. Reconcile the eligible count against the source-system extract before calculating any rate. Set identity, role, and reversal rules before rewards A behavior contract fails when one person appears under several IDs or several people share one ID. Choose the identity key used for eligibility and measurement. Email, phone, account ID, payment token, and household ID answer different questions; document merge priority plus the treatment of missing or changed identifiers. Identity checks are deterministic work. Flag duplicate source IDs, missing keys, impossible timestamp order, and one order attached to multiple customers. Reviewers decide ambiguous merges; code should enforce settled rules consistently rather than guess. Purchaser and value-creating participant may differ. A designated driver, event organizer, household shopper, or referrer can create the desired outcome without buying the rewarded item. If that is the intended behavior, store the role event separately instead of forcing it into an order definition. The distinction also appears in Designated Driver Rewards: Reward the Role, Not the Purchase . Define reversal timing too. A qualifying order refunded 10 days later should normally reverse both the event and its associated reward under a paid-purchase objective. Set a reporting delay matching the refund window, or publish provisional and settled measures separately. The classic failure: campaign selection uses account ID while reporting deduplicates by email. Treatment exposure and outcomes then operate at different units. Use one declared randomization and measurement unit, preserve its assignment, then reconcile exclusions and reversals against that unit. Test economics only after the event can be trusted Once the behavior is stable, assign eligible units to treatment and holdout before sending the reward. Analyze by original assignment, including treatment customers who never opened or redeemed. Removing them measures responders, not the effect of offering the program. Sample size must match the lift worth funding. Inputs are baseline rate, minimum detectable lift, two-sided significance level, desired power, treatment allocation, expected attrition, and any clustering by store or household. With an 18% baseline, a 3-percentage-point target lift, 5% significance, and 80% power, a standard two-proportion calculation needs roughly 2,700 customers per arm before attrition; 600 per arm will often leave that effect unresolved. Use a statistical power calculator or validated statistics package for the arithmetic. If randomization occurs by store, a customer-level calculation is insufficient because outcomes within stores may be correlated. Use a cluster-aware calculation based on the number of stores and an estimated intracluster correlation, or avoid a causal scale claim. Duration follows the behavior, not a fixed calendar rule. Required time equals recruitment period plus the complete outcome window plus any pull-forward window. For a 30-day repeat event with 21 days of recruitment and another 30 days needed to detect displaced purchases, the earliest settled read is 81 days after recruitment begins. Calculate incremental contribution per eligible customer as treatment contribution minus holdout contribution, then subtract incremental reward, communication, fulfillment, fraud, and variable operating costs. Measure cumulative contribution through the pull-forward window; merely looking for a later rate dip can miss changes in order value or margin. The classic failure: a team launches for six weeks because six weeks fits the planning cycle, then calls an immature cohort profitable. Prewrite the minimum detectable lift, sample requirement, maturity date, exclusions, profit equation, and stop rule. If the available audience cannot support the planned inference, narrow the claim or choose a larger commercially meaningful lift. Software comes after these definitions. When manual identity matching, reversals, or reward reconciliation become the constraint, use Loyalty Program Software: Buy for Requirements, Not Features to frame the purchase. Frequently asked questions What is the minimum viable repeat-behavior specification? Name the eligible population, eligibility timestamp, qualifying event, deadline, identity unit, accepted statuses, exclusions, and reversal rule. Add the numerator and matured denominator used for reporting. Should enrollment start the observation window? Only when enrollment itself creates eligibility for the behavior. For second-purchase measurement, anchor the window to the first qualifying purchase; otherwise enrollment timing changes the time available to repeat. Can a small customer loyalty program use a holdout? Yes, but a holdout does not guarantee a decisive result. Calculate sample needs first. A small audience may support only a larger detectable lift, pooled recruitment over more cohorts, or a directional estimate with explicit uncertainty. Should the first reward use points? Use points only when progress across purchases is part of the intended mechanism. A voucher or account credit requires fewer ledger states; points add accrual, expiry, reversal, liability, and fraud rules that must be reconciled. --- # Designated Driver Rewards: Reward the Role, Not the Purchase https://loyalflow.cc/blog/designated-driver-rewards-role-purchase The short version: Designated driver rewards test whether recognizing an occasion-enabling role produces more completed group visits or returns. Verify eligibility, randomize the offer, size the test from booking volume, then estimate absolute lift and cost per incremental visit. Key takeaways Reward the designated role with one fixed-value benefit per eligible booking. Randomize eligible bookings 50/50; self-selected participation cannot establish incrementality. Size the pilot from baseline behavior and minimum worthwhile lift, not a 6–8-week calendar. Track absolute attendance or return lift, confidence intervals, contribution margin, and total variable cost. Use deterministic booking, redemption, cancellation, and duplicate controls before considering points. Designated driver rewards value the group occasion Transaction-led loyalty rewards whoever spends. Group occasions work differently: a low-spend participant can remove a transport objection and make attendance possible for several other guests. That is a hypothesis worth testing, not proof that every designated driver creates incremental revenue. The smallest receipt may unlock the whole occasion. The useful unit is the eligible booking , not the driver's receipt. Record party size, venue, booking date, attendance, reward assignment, redemption, and subsequent completed bookings. Individual spend remains useful for margin analysis, but it cannot describe the whole occasion. A practical first rule: require a confirmed booking for at least 3–4 guests and nominate one participant before arrival. Give that participant a fixed $10–$20 benefit, such as a rideshare credit or nonalcoholic-drink allowance. Avoid a percentage discount; the reward recognizes a role rather than scaling with table spend. The trap is assigning economic value from one observed receipt. A driver ordering a $4 soft drink may have enabled the visit, or may simply have joined a visit that was already happening. Only a controlled comparison can separate those explanations. Test designated driver rewards with randomized bookings Do not compare bookings with a declared driver against bookings without one. Driver presence self-selects, so those groups may differ in party composition, travel distance, alcohol plans, or booking intent before the reward appears. Let chance choose before behavior can choose for you. Instead, determine eligibility first, then randomly assign eligible bookings 50/50 to treatment or control before revealing the offer. Stratify assignment by venue, weekday or weekend, and party-size band such as 3–4, 5–6, and 7+. This keeps obvious operating differences balanced without pretending matching removes every hidden difference. Choose one primary outcome before launch. Completed attendance is appropriate when the offer aims to prevent cancellation or no-show behavior. A completed return within 60 days is appropriate when the claim concerns retention, but only analyze bookings whose full 60-day observation window has elapsed. Size the test from the effect the economics require. As an illustration, detecting a return-rate change from 20% to 23% with 80% power and a 5% significance level needs roughly 2,900 bookings per arm. With 400 bookings per arm, the realistically detectable difference is closer to 8 percentage points. Exact requirements should come from a two-proportion power calculation using your baseline, minimum worthwhile lift, power, and significance level. Calendar length follows sample size. A venue network producing 1,000 eligible bookings weekly may finish quickly; one producing 100 weekly will not. Use 6–8 weeks as an operating window only when it supplies the required observations and covers representative weekdays, weekends, and pay cycles. The trap is calling matched or before-and-after results causal. Seasonality, local events, venue promotions, and booking mix can move outcomes without the reward. Random assignment gives the incrementality claim a procedure capable of testing it. Verify eligibility and cap cost before launch Eligibility control is deterministic. Require a unique booking_id , one named participant, confirmed attendance, one reward status, cancellation exclusion, and one redemption per booking. If a nonalcoholic purchase is required, capture the qualifying POS item instead of relying on staff memory. One eligible booking, one bounded benefit. Use a bounded redemption window, such as 24 hours before through 24 hours after the scheduled visit. The point is not that this window is universally superior. A bounded window limits unmatched records and late claims; widen it only when partner settlement timing requires more room. Set a fixed face value and hard issuance cap. A $20,000 budget with a maximum $20 reward permits no more than 1,000 issued rewards before partner fees, support costs, and taxes. Reserve those extra costs explicitly rather than discovering that the nominal voucher budget was not the pilot budget. Block duplicate IDs, cancelled bookings, reused reward codes, and ineligible venues immediately. Rate-based fraud stops need enough claims to interpret: two invalid claims among 40 submissions already equal 5%, but that estimate is unstable. Review an invalid-claim rate only after a predefined floor such as 200 claims, report its uncertainty, and keep immediate security blocks active regardless of sample size. The trap is treating self-declaration as verification. Shared screenshots, repeated contact details, and cancelled reservations can consume budget while producing no completed occasion. The minimum control set in Loyalty Program Fraud Prevention: Six Minimum Controls provides a useful extension. Make the decision from lift and margin Analyze the randomized groups as assigned, including treatment bookings that never redeem. Excluding non-redeemers selects customers after assignment and overstates the offer's effect. Report the treatment and control rates, absolute percentage-point lift, confidence interval, and number of assigned bookings in each arm. For attendance, divide attended bookings by all assigned eligible bookings. For 60-day return, divide groups with another completed booking by assigned groups whose observation window has fully elapsed. Define the group identifier before launch; changing from booker-level to participant-level identity after seeing results invites a favorable answer. Estimate incremental visits as treatment assignments multiplied by absolute lift. Then divide reward, partner, support, and payment costs by estimated incremental visits. If treatment attendance is 74% versus 70% across 1,000 assigned bookings, the point estimate is 40 incremental visits; uncertainty around the 4-point lift must accompany that estimate. Set the commercial threshold before launch. Break-even lift equals variable pilot cost per treatment assignment divided by contribution margin per incremental visit. If cost per assigned treatment booking is $4 and contribution margin per incremental visit is $50, break-even requires an 8-percentage-point lift before fixed setup costs. The trap is reporting 1,000 issued vouchers as success. Issuance measures distribution. The decision requires an estimated behavioral lift, uncertainty range, and contribution after reward costs. Survey approval can explain reactions, but repeat behavior determines the retention result; NPS vs Repeat Rate: Behavior Proves Retention covers that distinction. Source: www.marketingdive.com Frequently asked questions Should designated drivers earn loyalty points? Not in the first test. Use an immediate fixed benefit until the operator can identify the same participant across bookings and observe profitable repeat behavior over 30–60 days. Points add liability, expiration rules, support work, and another fraud surface. Who should fund the reward? Assign funding against measurable value. The venue may pay for incremental contribution, the brand for qualified participation, the booking platform for completed reservations, and the mobility partner for acquired rides. Document each party's maximum exposure, settlement field, refund rule, and dispute owner before launch. What customer data does the test require? Collect the minimum: booking_id , venue, visit date, party-size band, assignment, attendance, reward ID, issuance, redemption, and return outcome. If identifiable booking data passes to a brand, venue, booking platform, or mobility provider, disclose the fields received, purpose, retention period, and whether the recipient may use them for its own marketing. --- # Loyalty Program Software: Buy for Requirements, Not Features https://loyalflow.cc/blog/loyalty-program-software-requirements The short version: Loyalty program software should pass your mandatory workflows and financial controls at an acceptable three-year cost. Score observed execution, reconcile the pilot ledger, verify complete exports, then force integration failures before signing. Key takeaways Build a 100-point scorecard around 5–10 workflows tied to revenue, expected loss, frequency, or operator time. Require every mandatory financial and security control to pass; use weighted scores only to rank survivors. Compare 36-month cost, including implementation, usage, staff time , messaging, migration, and exit work. Reconcile the full pilot population when feasible; random spot checks cannot reliably expose rare or clustered defects. Test 3–5 end-to-end scenarios, including retries, refunds, suppression, exports, and forced failures. Score loyalty program software against real workflows Feature grids reward vendors for accumulating boxes. They do not show whether a refund reverses points correctly or whether support can repair an account without corrupting the ledger. Start with 5–10 workflows used by customers, operators, finance, and service teams. Only the load-bearing workflow earns the score. Assign 100 total points using company-specific exposure. Weight each workflow by transaction frequency, revenue affected, plausible financial loss, customer harm, and weekly operator time. One workable starting allocation is 20 points for earn and refund accuracy, 15 for redemption, 15 for liability reporting, 15 for service adjustments, 10 each for targeting, suppression, and exports, then 5 for permissions. Define the trigger, expected ledger entries, customer message, operator action, error state, and recovery path for every workflow. Score what the vendor demonstrates in your scenario, not a claim that the feature is supported. Require every mandatory control to pass; rank the remaining vendors by weighted score and three-year cost rather than using an arbitrary aggregate cutoff. The classic failure: an 80-feature checklist lets cosmetic features offset a broken reversal flow. They are not substitutes. A theme editor cannot repair duplicated points after an order refund. Calculate three-year cost, not subscription price Model 36 months because implementation work, usage tiers, migrations, and recurring administration rarely appear in the headline price. Include subscription fees, transaction or member charges, email and SMS, integration work, premium support, sandbox access, data migration, internal administration, and exit assistance. The monthly price is only what shows above ground. Run low, expected, and high cases for active members, monthly transactions, messages, and redemptions. Price internal work at a loaded hourly cost. Ten operator hours per week equals roughly 1,560 hours over three years, before holidays or volume growth. Keep platform cost separate from reward economics. A 2% earn rate on $5 million of eligible sales issues $100,000 of face value before redemption and expiration assumptions. Contract treatment of unused points, vendor-funded rewards, expiration, and breakage affects cash planning and accounting. Price the pilot from required vendor hours, environments, integrations, migration volume, and support coverage. Set its duration from the cycles the test must observe. A 30–45 day window can fit routine purchase, refund, messaging, and reporting tests, but it is too short when acceptance depends on quarterly tiers or longer expiration rules. The classic failure: a $2,000 monthly platform appears cheap while requiring weekly CSV repair, custom middleware, and paid escalation. Put every recurring manual task into the cost model before comparing licenses. Test the ledger and exports across the full pilot population Points require ledger discipline. Opening balance plus earns, adjustments, and reversals minus redemptions and expirations must equal closing balance for every account and for the aggregate program. Arithmetic, duplicate detection, referential integrity, and row-count comparison belong in queries or deterministic rules, not operator judgement. Inspect the whole net; one missed knot still leaks value. Test partial and full refunds, negative balances, delayed events, duplicate delivery, retries, expired points, reinstatement, account merges, and manual adjustments. Every mutation needs a timestamp, reason, source event, actor, and immutable identifier. Finance should reproduce period-end balances without a vendor-built private report. Reconcile every pilot ledger event when volume permits. Compare source-event counts, ledger-entry counts, unique identifiers, control totals, and closing balances. Test every forced retry, timeout, refund, import, and merge because duplicate defects cluster around those mechanisms; a random sample may never touch them. If full-population reconciliation is impractical, define the defect rate the sample should detect and the required confidence before sampling. Under independent random sampling, detecting a 0.1% defect with about 95% probability requires roughly 2,995 observations. Then stratify by event type, integration, day, refund state, latency, and retry status; report exceptions against the denominator in each stratum. Demand exports for customers, accounts, transactions, rewards, consent, tier history, expiration dates, and source references. Verify fields including customer_id , transaction_id , source_order_id , timestamps, currency, and status. Compare complete export row counts and control totals with the application, then reconstruct balances in your warehouse or spreadsheet. The classic failure: “100% of sampled events reconciled” sounds conclusive without sample size, selection method, population coverage, or a detectable-defect target. A clean sample of 100 has only about a 9.5% chance of finding an independently distributed 0.1% defect. Use the control framework in Loyalty Points Liability: Build Controls Before Campaigns . Force integrations and recovery paths to fail An integration badge proves that two systems exchanged something. Test 3–5 production-like scenarios covering enrollment, purchase and earn, redemption, refund or cancellation, and lifecycle suppression. Record expected events in every system, acceptable latency, retry behavior, failure ownership, and operator recovery. Break it before customers do—and verify the reset. Financial events should reconcile exactly. Set marketing latency from the business deadline: a suppression must arrive before the next scheduled send, not within an arbitrary vendor benchmark. Force one timeout, one duplicate delivery, one missing identifier, and one permission failure; confirm idempotency, alerting, recovery, and audit history. Use representative products, customer states, returns, consent records, and transaction volumes. Synthetic data covers destructive edge cases safely. If real customer data enters a vendor environment, document the disclosed identifiers, purchase history, balances, and consent fields, plus storage location, access, retention, and deletion terms. Define acceptance before configuration starts: all mandatory controls pass, all pilot ledger events reconcile or meet the documented sampling design, required export fields are complete, forced failures recover correctly, and operator time stays below your cost-model ceiling. Routine CSV patches count as failed automation. The classic failure: a vendor-led happy-path demo proves enrollment, then launch exposes missing refund reversals and promotions sent during service disputes. Define that second control with Lifecycle Suppression Rules: Stop Marketing Through Service Failures . Frequently asked questions How many vendors should make the shortlist? Score 4–6 vendors from documented evidence, invite 2–3 to scripted workflow demonstrations, then pilot the highest-ranked candidate that passes every mandatory control. Pilot a second vendor only when scores and three-year costs remain materially close. Which conditions increase migration risk? Missing transaction history, unstable customer identifiers, undocumented expiration rules, duplicate accounts, and unreconciled balances increase cutover uncertainty. Profile each defect as both a count and percentage of its relevant population before migration. When should a company build instead of buy? Build when distinctive loyalty logic creates enough economic value to justify continuing ownership of the ledger, security, support, compliance, and migrations. Buying transfers baseline product maintenance to a vendor; custom integrations and company-specific rules remain your responsibility. --- # Costco Membership Model: Renewal as Pricing Governance https://loyalflow.cc/blog/costco-membership-model-pricing-governance The short version: The Costco membership model is distinctive because membership renewal governs merchandise pricing. The fee creates recurring contribution; the risk of losing that fee constrains markups and protects the customer bargain. Key takeaways Treat renewal as a pricing constraint, not a marketing score. Separate membership contribution from merchandise contribution. Calculate allowable CAC from contribution and renewal assumptions. Test light, base, and heavy members before approving benefits. Copy Costco only when customers can verify value repeatedly. The Costco membership model governs pricing Most paid loyalty programs sell a bundle of benefits. Costco sells access to a retail system whose credibility depends on disciplined merchandise pricing. That difference matters: the membership fee does not merely fund perks; it gives the operator a recurring profit pool worth protecting. A retailer funded mainly through transaction margin can improve short-term profit by raising prices. Costco faces a counterweight. Higher merchandise margin may help this quarter, but obvious price deterioration weakens the argument for paying the next annual fee. Renewal therefore acts as pricing governance. Buyers still negotiate products, merchants still manage categories, and operators still pursue transaction contribution. The membership model places a ceiling on extraction because members can reassess the entire bargain once each year. This is the useful distinction from a generic paid-membership model. A program can charge $60, distribute $60 of shipping or credits, then hope incremental purchases cover the difference. Costco’s mechanism runs in the opposite direction: recurring fee contribution supports tighter merchandise economics, while those economics defend renewal. Operators can test this logic with one question: if merchandise margin rose by 2 percentage points, would members notice enough deterioration to reduce renewal? If nobody would notice, the fee probably is not governing pricing. The business merely has a subscription attached to ordinary retail. The classic failure: using membership revenue as permission to stack more margin on top. Customers then pay once for access and again through uncompetitive prices. Renewal becomes dependent on inertia rather than demonstrated value. Separate fee contribution from merchandise contribution Do not judge the model using membership revenue alone. Build two contribution lines. Membership contribution equals fee revenue minus benefit cost, payment fees, incremental service cost, and expected refunds. Merchandise contribution equals member gross profit minus fulfillment, returns, and other variable transaction costs. Two profit pools, one clearer decision. Suppose the annual fee is $60. Payment, service, and included-benefit costs total $18. Membership contribution is $42 before any merchandise activity. If the representative member generates another $35 of annual merchandise contribution, total annual contribution is $77. Now test a proposed price concession. Reducing merchandise contribution by $10 may be rational if the stronger value proposition protects more than $10 of expected fee contribution. That is the governance trade: sacrifice some transaction margin only when renewal economics plausibly repay it. Do not hide the trade inside blended gross margin. Review fee contribution, merchandise contribution, realized member savings, and renewal by signup cohort. Monthly review still matters for an annual plan because first purchase, second purchase, and first visible saving should usually occur within 30–45 days. Set the initial 30-day activation threshold as a launch hypothesis, not an industry benchmark. For example: at least 60% of new members must use the core value mechanism within 30 days. After two or three cohorts, replace that threshold with the activation level actually associated with profitable renewal. The classic failure: reporting the $60 fee as pure margin while benefits, support, refunds, and price concessions sit in other budgets. The program looks profitable because its costs have been scattered across the company. Use renewal math to set allowable CAC A universal renewal target is useless without contribution and acquisition cost. Calculate the renewal rate your model requires instead of borrowing a benchmark from a larger operator with different frequency, margins, and acquisition channels. CAC can stretch only as far as renewal pays. For a simple annual model, estimated lifetime contribution equals first-year contribution plus annual contribution multiplied by renewal probability divided by one minus renewal probability. This assumes constant contribution, constant renewal, and no discount rate, so use it as a screening calculation rather than a forecast. Using $77 of annual contribution and a 60% renewal assumption, expected future contribution is $77 multiplied by 0.60 divided by 0.40, or $115.50. Add the first year and estimated lifetime contribution is $192.50. A $120 CAC leaves only $72.50 before overhead, forecasting error, and capital cost; a $40 CAC leaves substantially more room. Run the same calculation at 40%, 60%, and 75% renewal. These are scenarios, not recommended targets. If the acquisition case works only at 75%, reject it until observed cohorts support that assumption. Then reverse the formula. Start with observed annual contribution, subtract the profit buffer required by the business, and treat the remainder as maximum CAC. Do not increase CAC because a blended renewal number looks healthy; calculate it separately for paid search, referrals, store conversion, partnerships, and promotional cohorts. The classic failure: assuming a high renewal rate makes any CAC affordable. Weak annual contribution can still produce poor lifetime economics, while aggressive first-year discounts can attract cohorts that disappear at full-price renewal. Run a light, base, and heavy-member worksheet A launch decision needs three scenarios. Use the same fields for each: annual fee, eligible spend, price concession, merchandise contribution, benefit cost, service cost, renewal assumption, and CAC. A spreadsheet with three rows is enough. One model, tested at three levels of appetite. Consider a $60 membership offering a 5% price advantage on eligible purchases. A light member spending $300 receives $15 of visible value. That member may be profitable but unlikely to renew because the fee remains unrecovered. A base member spending $1,500 receives $75 of visible value. If ordinary merchandise contribution before the concession is 20%, the 5% concession leaves $225 of merchandise contribution rather than $300. Add the fee, subtract $18 of membership costs, and annual contribution becomes $267 before CAC and overhead. A heavy member spending $5,000 receives $250 of visible value. The member may still produce strong contribution if the 5% concession applies to purchases with sufficient underlying margin. If it applies to low-margin categories, the same customer can become loss-making despite high revenue. The decision rule is blunt. Light members need enough early proof to renew. Base members must repay acquisition and operating costs. Heavy members must remain contribution-positive without manual exclusions. If one benefit cannot satisfy all three, narrow eligible categories, cap usage transparently, or reject the plan. The classic failure: designing around the average member. Averages conceal light users who receive no credible value and heavy users who consume every subsidy. Model both tails before collecting a fee. This mechanism also explains why categories with long replacement cycles should resist warehouse imitation; infrequent-purchase loyalty usually needs a recurring service or access habit instead of transaction discounts . Frequently asked questions What renewal rate should a Costco-style membership target? No universal rate. Calculate the minimum rate required for lifetime contribution to cover CAC, overhead allocation, and a forecast-error buffer. Treat early rates as hypotheses until several renewal cohorts mature. How quickly should members recover the fee? Design for the representative member to see credible progress within 30–45 days and recover the fee during normal annual purchasing. Requiring abnormal spend makes the value proposition promotional rather than durable. When should an operator copy the Costco membership model? Copy it when customers purchase often, compare prices easily, and can verify recurring savings. Skip it when value depends on obscure calculations, rare purchases, or benefits whose variable cost rises faster than member contribution. --- # Loyalty Program Fraud Prevention: Six Minimum Controls https://loyalflow.cc/blog/loyalty-program-fraud-prevention The short version: Loyalty program fraud prevention starts with six controls, not a scoring platform: stronger authentication, risk-based redemption friction, referral limits, complete event logging, a transactional points ledger, and an operable review queue. Key takeaways Build a 90-day loss register before choosing fraud tools. Apply step-up verification to sensitive changes and valuable redemptions. Keep referral rewards pending through merchant-specific refund and dispute exposure. Process points through an atomic, idempotent, append-only ledger. Start with 3–5 explicit rules; report alert counts, precision, false positives, and loss. Start loyalty program fraud prevention with loss paths Map four loss paths first: account takeover, referral abuse, unauthorized points use, and staff adjustments. For each path, record incident count, attempted value, confirmed loss, recovered value, and investigation time over the previous 90 days. No history means start logging now, not inventing a risk score. Measure where value escapes before buying a better alarm. Compare frequency, value, and recoverability separately. One hundred $20 referral claims create different economics from one $2,000 account takeover. An unshipped order may be stopped; a transferred gift card may become unrecoverable within minutes. Logging enables attribution, reconciliation, and review. Capture customer_id , event time, session or device identifier, IP address, authentication result, profile changes, points before and after, order reference, actor, and reason code. Retain history through the longest applicable refund, dispute, and appeal period; 180–365 days is a provisional operating range, subject to legal and privacy requirements. Failure mode: unreconciled vendor scores. A platform produces alerts, but the team cannot connect them to confirmed loss, recovered value, customer harm, or points liability. Define the loss register and event schema before buying another alert feed. Control account takeover without challenging every visit Require MFA immediately for administrators and staff with adjustment rights. For customers, use step-up verification after combinations such as a new device plus profile change, a password reset followed by redemption, or a high-value portable reward claim. Small rewards pass lightly; portable value earns a harder check. A $50 reward threshold can be a provisional starting point, not a universal benchmark. Add an absolute ceiling, then adjust against normal order value, confirmed losses, and queue capacity. Percentage-of-annual-earn rules need special handling for accounts under 90 days old or with sparse history; those accounts should use the absolute threshold and account-age rules instead. Block known breached passwords during creation and reset. Invalidate existing sessions after password, email, phone, or MFA changes. Send immediate alerts with a recovery route that does not depend only on the potentially compromised email address or phone number. Consider a 24–72-hour redemption hold after sensitive profile changes when rewards are portable or hard to recover. Make that range provisional, measure abandoned legitimate redemptions, and allow documented manual verification. The hold buys response time; it does not establish fraud. Redemption friction should rise with recoverability and value. A $5 discount attached to an existing order may need no extra step. A $500 gift card, points transfer, or shipment to a new address warrants fresh authentication and may warrant approval. Failure mode: friction in the wrong place. MFA on every visit frustrates legitimate members while weak profile-change and redemption controls still let an attacker replace the recovery channel and cash out. Put referrals and points under transaction controls Referral programs need hard limits by customer and time window, with device, payment token, and delivery address used as review signals. Five rewarded referrals per customer per 30 days is a provisional test limit. Compare it with legitimate advocate behavior before enforcing it broadly; households, offices, and apartment buildings legitimately share attributes. Every points movement lands once—and stays accountable. Keep referral value pending until the qualifying transaction clears the merchant’s cancellation and refund deadlines. Check the actual processor and card-network dispute exposure rather than assuming 45–90 days covers it. If waiting through the full dispute period is commercially unacceptable, release after the return window while retaining disclosed reversal authority or funding a reserve for later disputes. Points need a ledger, not a mutable balance field. Record every earn, redemption, expiration, adjustment, and reversal as a new transaction. Store the original transaction reference on reversals and enforce an idempotency key so retries cannot redeem twice. Redemption must use a database transaction or conditional write that checks and deducts the available balance atomically. Reject unauthorized negative balances. This is deterministic transaction integrity, not a reviewer decision or fraud-model task. Use count and value velocity rules together. Provisional triggers might include more than 3 redemptions in 10 minutes, over $250 in reward value within 24 hours, or redemption from 2 new devices within 7 days. These conditions trigger review or step-up verification; they do not confirm fraud. Separate requester and approver for manual adjustments above a provisional $100–$250 threshold, calibrated to normal order and reward value. Log actor, reason code, linked case, previous balance, and resulting balance. Failure mode: trusting the displayed balance. Concurrent requests can both spend the same points unless deduction is atomic and idempotent. Missing actor and reason records also let staff adjustments become unexplained balance changes. Operate rules and review before adding scoring Start with 3–5 reproducible rules. Examples include new device plus profile change within 24 hours, password reset plus redemption within 60 minutes, or several accounts sharing a payment token and redeeming to one address. Three failed verifications followed by success should trigger step-up verification, never a fraud conclusion. Rules detect explicit conditions; arithmetic reconciles the ledger. Scoring becomes useful only after resolved cases provide labels for comparison. It cannot repair missing events, duplicate transactions, or broken balance calculations. Every alert needs evidence, an owner, and a deadline. Show triggering events, linked accounts, reward value, transaction history, authentication history, and customer contacts. Record fixed outcomes: confirmed fraud, legitimate, insufficient evidence, customer error, or policy abuse. Target review within 4 business hours for held, irreversible rewards and no later than 1 business day where staffing allows. Set appeal targets of 2–5 business days. These are service targets to test against queue volume, not claims that every team can meet them immediately. Publish counts beside every rate. Use precision = confirmed fraud alerts / resolved alerts . Use false-positive share = legitimate alerts / resolved alerts . Calculate review rate as reviewed eligible events divided by all eligible events, and recovery rate as recovered confirmed loss divided by confirmed recoverable loss. A customer false-positive rate requires a harder denominator: legitimate eligible events incorrectly held divided by all legitimate eligible events. That denominator may require later outcome matching. Never call every declined or held redemption prevented loss. Review each rule after at least one complete 30-day operating cycle and a predefined volume floor, such as 50 resolved alerts. Report small samples as counts, not confident rates. Keep a low-precision rule when it catches material, unrecoverable loss; remove or narrow rules that create work without actionable decisions. Failure mode: alert volume without denominators. Forty rules can fill a queue while revealing nothing about eligible transaction volume, customer impact, or confirmed loss. Start small, assign ownership, then tune using resolved cases and loss value. Ledger controls also protect reward accounting. Pair this fraud plan with Loyalty Points Liability: Build Controls Before Campaigns , then use Auditing Loyalty Data With AI: Most of the Job Is Not AI when checking event and balance integrity. Frequently asked questions When should customers face MFA? Use step-up MFA for sensitive profile changes, new recovery channels, portable rewards, and high-value redemptions. Avoid challenging every login unless measured account-takeover risk justifies that friction. Staff and administrator MFA should be mandatory immediately. How long should referral rewards remain pending? Use the qualifying purchase’s actual cancellation, refund, and dispute deadlines. If holding through the full dispute period would damage the referral offer, release after returns while retaining disclosed reversal authority or maintaining a reserve for later losses. How should a small team review suspicious redemptions? Begin with 3–5 rules, one daily owner, one backup, and fixed outcome codes. Prioritize irreversible rewards and provisional values above $50–$100. Review counts weekly; assess rates only after a defined cycle and sufficient resolved volume. What customer data does an external fraud vendor receive? Depending on the integration, the vendor may receive customer identifiers, transaction history, payment tokens, addresses, device identifiers, and network data. Document that disclosure, restrict fields to operational need, set retention terms, and confirm customer-facing privacy notices cover the transfer. --- # Shopify Loyalty Program Setup: A 30-Day Implementation Plan https://loyalflow.cc/blog/shopify-loyalty-program-setup Shopify loyalty program setup fails in the plumbing: rewards disappear from the cart, refunded orders keep their points, guest customers create duplicate accounts, and reports cannot separate members from everyone else. A clever points model cannot rescue broken implementation. Use 30 days to configure one earn rule, one reward, four customer-facing surfaces, and one controlled measurement plan. Do not add tiers, referrals, birthday rewards, or bonus campaigns until the basic transaction works from order creation through refund. Define the Shopify Data Flow Before Choosing Features Map six events during days 1–7: customer enrollment, eligible order, points approval, reward issuance, reward redemption, and order refund. For each event, identify the Shopify customer ID, order ID, timestamp, value, and status available in your loyalty app export. One durable ID keeps the loyalty ledger honest. Use Shopify customer IDs as the primary customer key. Email addresses change, guest checkouts create duplicates, and phone numbers arrive in inconsistent formats. Test whether the app merges a guest order after account creation or leaves two balances requiring manual repair. Set points to pending until the refund window closes. If most refunds arrive within 14 days, use 14 days as the starting assumption; if apparel returns remain open for 30 days, use 30. The setting should follow actual return data, not the app default. Write explicit rules for full refunds, partial refunds, cancellations, edited orders, gift cards, shipping, tax, discounts, and subscription renewals. A defensible default earns value on net eligible product spend after discounts, excluding tax, shipping, and gift-card purchases. Implementation trap: points become available when an order is placed, then get spent before the original order is refunded. The store loses the reward and the original revenue. Pending points plus automatic reversal closes that hole. This week’s test: place one order containing two products, a discount, tax, and shipping. Partially refund one item. The resulting balance must match the written rule without manual adjustment. Configure the Smallest Shopify Loyalty Program Setup During days 8–14, enable one purchase earn rule and one fixed-value reward . Customers should be able to explain both in two sentences. Disable tiers, social actions, referrals, multipliers, and automatic birthday grants. Generosity works better when the math stays balanced. Calculate the reward rate from margin tolerance rather than copying a benchmark. Use: maximum issued reward rate = allowable contribution-margin reduction divided by expected redemption rate . Treat redemption as an assumption until the store has its own data. Example: the store can tolerate a 2% reduction in contribution margin per eligible dollar. At an assumed 50% redemption rate, the maximum issued reward rate is 4%. At 25% redemption, the same model permits 8%, but building economics around high breakage is reckless; redemption can rise once customers understand the program. Check attainability against average order value and reorder timing. A $5 reward requiring $100 of cumulative eligible spend may suit a store with a $65 average order value and a 45-day reorder cycle. It will feel remote for a store with a $25 average order value and two purchases per year. Choose a threshold reachable after one or two normal orders without distorting basket behavior. Show the calculation in dollars, even if the app displays points: reward value divided by required spend equals the issued reward rate. Configuration trap: using 1,000 points for a $5 reward because large balances look exciting. Customers still receive $5; support now has to explain conversion math. Use the lowest point scale the app permits cleanly. Go forward only if reward cost remains inside the predeclared margin floor under low, expected, and high redemption scenarios. For a new program, 20%, 40%, and 60% are scenario inputs, not industry benchmarks. Place Rewards Inside Shopify’s Buying Surfaces During days 15–21, test four surfaces on mobile and desktop: customer account, product or collection page, cart, and post-purchase message. A floating launcher can support these placements; it cannot replace them. A reward unseen at checkout is barely a reward. The customer account should show approved balance, pending balance, available rewards, expiration terms, and transaction history. The cart should show dollar value rather than points alone: “$5 reward available” or “Spend $28 more to reach a $5 reward.” Confirm how the reward enters checkout. Test discount-code conflicts, automatic discounts, subscription products, sale items, minimum baskets, accelerated checkout, and multiple currencies where applicable. If Shopify plan or checkout restrictions block a placement, move the message upstream into the cart rather than promising unavailable checkout behavior. Send a post-purchase message after points become approved, not merely when the order is placed. Include eligible spend, points earned, pending or approved status, current reward value, and a direct route back to the store. Storefront trap: the program page promises rewards that disappear when a customer uses Shop Pay, a subscription item, or an existing discount. Test the actual checkout combinations responsible for most revenue, not only a clean test order. Before launch, complete at least 10 end-to-end transactions covering guest checkout, account login, discount stacking, cancellation, full refund, partial refund, reward redemption, failed payment, subscription renewal, and mobile checkout. Every failure needs either a configuration fix or clearly displayed restriction. Launch With Formulas, Comparison Groups, and Stop Gates Use days 22–30 for a controlled launch. Split one eligible customer segment into exposed and holdout groups where tooling permits. Keep acquisition channel, first-order month, geography, and initial order value reasonably balanced; comparing volunteers with non-members will overstate results because frequent buyers enroll more readily. Launch it like an experiment, not a leap of faith. Define activation rate as customers who earn approved value or redeem a reward divided by enrolled customers. Define redemption rate as reward value redeemed divided by reward value issued. Define repeat-purchase rate as customers placing another eligible order within the chosen window divided by first-time customers in that cohort. Calculate cohort lift as exposed-group repeat-purchase rate minus holdout-group repeat-purchase rate. Calculate contribution margin per customer as net revenue minus product cost, discounts, reward cost, payment fees, fulfillment, shipping subsidy, and variable app cost, divided by customers. Set gates before launch. One practical pilot rule is no broad expansion before each arm has at least 200 eligible customers; this is an operating threshold, not proof of statistical significance. Report the observed difference with a confidence interval, then wait for more data if the interval still includes material loss and material gain. Pause immediately if reward errors affect more than 1% of tested transactions, balances fail to reverse after refunds, or contribution margin per customer falls below the declared floor. Review direction after 30–60 days for short reorder cycles and 90–120 days for slower categories. Measurement trap: calling enrollment growth retention. Ten thousand members with unchanged repeat purchases and lower contribution margin represent a discount system, not a loyalty asset. After the pilot, remove placements nobody sees, messages nobody acts on, and rules support cannot explain. If the core transaction works but the reward structure still underperforms, use Points, Tiers, or Cashback to decide whether the model—not the Shopify setup—needs changing. --- # Auditing Loyalty Data With AI: Most of the Job Is Not AI https://loyalflow.cc/blog/ai-loyalty-data-audit The short version: duplicate customer identities are a defect worth auditing early, because they distort segments quietly and the damage compounds. This is the shape of that job — not a runbook. The model comes last: normalisation and deterministic matching resolve whatever they can, blocking is what makes the problem computable at all, and the model judges only the ambiguous remainder. The step most write-ups omit is measuring the result against a slice you checked exhaustively, and that is the step that decides whether any of it worked. Bad loyalty data does not announce itself Some data problems fail loudly: a broken export, a schema change, a job that errors. Those get fixed, because something is visibly wrong. Identity fragmentation fails quietly. You still get five segments. Champions still have the highest scores. The report renders. Nothing tells you that your best customer exists as three profiles — a work email, a personal email, and a guest checkout — each looking like an ordinary two-order buyer, none looking like the ten-order customer they actually are. The consequence is not a wrong chart. It is a reward budget aimed at people who were never going to leave, while the fragmented customer sits in a segment that gets a win-back discount they did not need. The RFM spreadsheet method recommends inspecting twenty records before scoring: the five highest monetary values, five highest frequencies, five longest recencies, and five at random. That is a sound smoke test and it is what it says it is — a check for gross errors before you trust the file. It cannot tell you how many duplicates are in eighty thousand profiles, and it does not claim to. Start with normalisation, not with a model Before anything clever, flatten the obvious variation: Emails: lowercase, strip dots and +tags where the provider ignores them, trim whitespace. Phones: strip formatting, normalise to E.164, handle the local-vs-international prefix for your markets. Names: case-fold, strip titles and punctuation, normalise accents. Addresses: standardise to a postal format if you have a library for your country. Then match deterministically on the identifiers that are supposed to be unique: normalised email, normalised phone, payment token, loyalty card number. Where two profiles agree on one of those, you usually have a duplicate and you did not need a model to say so. How much this catches depends entirely on your data — how many customers use one email across purchases, whether guest checkout captures a phone, whether your payment provider exposes a stable token. Measure it on your own file rather than trusting anyone's figure, including this one. The point is that this step is cheap and its output is auditable, so it should absorb everything it can before you spend money on inference. Blocking: why you cannot just ask the model Here is the constraint that decides the whole design. Eighty thousand profiles produce roughly 3.2 billion possible pairs. You cannot send that to anything. Even at a hundredth of a cent per comparison it is an unaffordable job, and most of those pairs are two people who share nothing. So you generate candidates first. Group profiles into blocks that share something cheap and discriminating, and only compare within blocks: Same phone suffix (last six digits) Same first three characters of surname plus year of first order Same normalised street number plus postcode outward code Choose keys that are discriminating . A key that groups everyone in a dense urban postcode produces a block of thousands and has not reduced anything — check the size distribution of your blocks before trusting a key, and drop any that produce a long tail of huge ones. Use several keys, not one: a single key misses every duplicate where that field is wrong or absent, which is exactly the population you are hunting. Their union is what shrinks the problem, and by how much depends on your data — measure it. This step is pure code. No model. It is also the step most "use AI to dedupe your CRM" advice omits, which is how you end up with a workflow that cannot run on real data. Where the model actually earns its place You now have a few thousand candidate pairs. Compute comparison features for each one — string distance on the name, whether postcodes match, days between first orders, overlap in purchased categories, whether the domains differ — and let a scoring step rank them. For clearly-similar and clearly-different pairs, a threshold on those features is enough. The model is for the middle: two records that share a surname and a city but differ in everything else, a business name against a personal name at the same address, a transliterated name spelled two ways. Two rules that matter more than the model choice: Explanations must come from the computed features, not from the model's own account of itself. "Same postcode, card last four matches, names differ by one edit" is checkable. A free-text rationale the model composed can describe evidence that is not there. The output is a queue, never an action. Nothing merges automatically. This is not general caution — a wrong merge is unusually hard to reverse, because the record you would need to undo it is the one that got absorbed. Measure it against a slice you checked completely This is the step that gets skipped, and skipping it is how you end up confident and wrong. The intuitive check is to sample: take twenty pairs the model flagged and twenty records it passed, and inspect them. The first half is fine — it estimates precision, roughly, on a small sample. The second half does not work , and it is worth seeing why. Suppose duplicates are 2% of your records and the tool misses half of them. Then about 1% of the records it passed are misses. Sample twenty of them and your chance of catching even one is about 18%. You will almost certainly see nothing, and "I checked and found no misses" is precisely the wrong conclusion to draw from a test that fails to fire four times out of five. Do this instead. Take a narrow slice you can examine exhaustively and find every duplicate in it by hand. That gives you a labelled set with a known denominator, which is what the sample lacked. Run the tool over the same slice and compare: what fraction of real duplicates it found, and what fraction of its claims were right. Define the slice carefully, because the obvious choices leak. A surname slice misses the duplicate who married and changed name; a one-month order slice misses the same person's orders in other months. Search each in-slice profile against the whole customer table, not just against the slice, and write down the rule you used to call something a match before you start — otherwise you are labelling to match what the tool found. A small complete slice beats a large random sample, because it is the denominator that makes the numbers mean anything. This also fixes the reverse error: precision of 45% sounds poor, but if duplicates are 1% of pairs it means the tool concentrated the problem forty-five-fold, and that may be an excellent queue to work through. Precision without a base rate is not interpretable in either direction. What you send, and what you keep The article you are reading recommends sending customer records to a model. Say plainly what that means: unless you are running something locally, those records leave your infrastructure and land with a third party under their retention and training terms. Minimum discipline: send the computed comparison features rather than raw records wherever the judgement allows it; drop every field the decision does not need; check whether your provider trains on submitted data and turn that off; confirm the arrangement is covered by your processor agreements before anything is exported, and that your retention and deletion terms cover it. Two traps worth naming. Hashing is not anonymisation — an email or phone hash is pseudonymous and trivially reversible by enumeration, so a hashed identifier still carries the obligations of the original. And payment tokens are payment data : if you use a token prefix as a blocking key, that key stays inside your own infrastructure, and only the resulting comparison feature — matched or not — travels. This is a summary, not a compliance assessment. A deduplication project that creates a disclosure problem has not improved your data. Two jobs where AI is the wrong tool Worth stating, because both get sold as AI use cases. Points ledger reconciliation is accounting, not inference. Whether an account's balance rolls forward correctly, whether an adjustment carries a reason code, whether an approval exists — these are deterministic checks with exact answers, and a model can only make them less certain. Note also that individual earn and burn events are not supposed to balance, and negative balances are sometimes legitimate after a reversal, so the rules must encode your actual program terms. See points liability controls for the framework this sits inside. A model is useful at one edge only: classifying free-text adjustment notes after the deterministic checks have run, and even then with a defined taxonomy and an "unclear" option. Segment drift is a profiling job. If Champions shrank this month, compare field distributions, null rates, and volumes against previous periods, and check your deployment log. That is measurement. A model asked to explain the shift will produce a plausible narrative, which is worse than no answer because it is persuasive. What to run this month Normalise email, phone and name on a copy of your customer table. Count how many exact duplicates appear on each identifier alone. This costs an afternoon and may be most of your answer. Build two or three blocking keys and count the candidate pairs they produce. If the number is not in the thousands, adjust the keys before going further. Pick a slice you can check exhaustively and label it by hand. Do this before running anything, so the tool cannot influence what you count as a duplicate. Score the candidates, measure against the labelled slice, and only then decide whether the queue is worth working. After merging confirmed duplicates, recompute your retention inputs from scratch — frequency, contribution margin, cohort retention — rather than adjusting the old figures. The retention calculator recomputes lifetime value from those inputs, so feed it the corrected ones and compare. Merging generally raises per-customer frequency, but the direction of the final number depends on which inputs moved. Keep the margin-based definition the whole way through. Key takeaways Identity fragmentation is quiet and survives into the budget, which is what makes it worth auditing early. Normalisation and deterministic matching come before any model. How much they resolve depends on your data — measure it before buying anything. Blocking is what makes deduplication computable — eighty thousand profiles are 3.2 billion pairs. Judge the result against a slice you checked exhaustively. Sampling the records a tool passed cannot measure what it missed. Precision means nothing without a base rate; 45% may be excellent or useless depending on the denominator. Keep it read-only, send the minimum, and confirm your privacy policy already covers it. --- # RFM Segmentation Spreadsheet: Build Five Usable Segments https://loyalflow.cc/blog/rfm-segmentation-spreadsheet An RFM segmentation spreadsheet can replace weeks of analytics work. Export 12 months of orders, calculate three scores, map every customer into five exclusive segments, then refresh monthly. That is enough for most teams to make better retention decisions without a warehouse or predictive model. Do not optimize for analytical elegance. Optimize for a file one operator can refresh in under an hour and use without interpretation meetings. If a segment does not change treatment, delete it. Build the RFM Segmentation Spreadsheet Create an Orders sheet with customer ID in column A, order date in B, net revenue after refunds in C, and order status in D. Use completed orders only. Exclude tax, shipping, cancellations, gift-card purchases, and unidentifiable guest orders where possible. Create a Customers sheet with one row per stable customer ID. Put a fixed scoring date in B1. Assuming the customer ID is in A2, calculate recency with =B$1-MAXIFS(Orders!B:B,Orders!A:A,A2,Orders!D:D,"Completed") . Calculate frequency with =COUNTIFS(Orders!A:A,A2,Orders!D:D,"Completed") . Calculate monetary value with =SUMIFS(Orders!C:C,Orders!A:A,A2,Orders!D:D,"Completed") . Keep the fixed scoring date instead of TODAY(). Otherwise, two people opening the same file on different dates can produce different segments. Use a 12-month order window when normal repurchase occurs within 90 days. Test 18 or 24 months only when the median interval between orders exceeds six months. Data trap: one refunded wholesale order can create a false Champion. Duplicate profiles can turn a five-order buyer into three weak buyers. Before scoring, inspect at least 20 customer records: the five highest monetary values, five highest frequencies, five longest recencies, and five random rows. Score Quintiles Without Splitting Ties Store the 20th, 40th, 60th, and 80th percentiles for each metric in visible cells. For frequency in column C, the thresholds are =PERCENTILE.INC(C:C,0.2) , then 0.4, 0.6, and 0.8. Repeat for monetary value and recency. Equal customers stay together, even when the buckets come out uneven. For frequency or monetary value, assign scores with =1+(C2>$H$2)+(C2>$H$3)+(C2>$H$4)+(C2>$H$5) , where H2:H5 contains the four thresholds. For recency, where lower is better, use =5-(B2>$G$2)-(B2>$G$3)-(B2>$G$4)-(B2>$G$5) . The strict greater-than comparisons keep equal values together. That matters when thousands of customers have exactly one order. Quintiles may become uneven; accept that. Forcing exactly 20% into each bucket splits identical customers for no operational reason. If you would rather not build the sheet, the RFM segmentation tool runs these same formulas on an orders CSV in your browser: the same PERCENTILE.INC quintiles, the same strict comparisons, and the same five segments. Nothing is uploaded. Keep R, F, and M in separate columns. Do not sum them. A 155 customer bought recently but rarely; a 551 customer bought often and spent heavily but has gone quiet. Both total 11. They need opposite treatment. Scoring trap: turning all 125 combinations into campaigns. Preserve the three-digit score for analysis, then collapse it for execution. Below roughly 500 identifiable customers, start with three score bands rather than five because small percentile shifts can move too many customers between buckets. Map Every Score Into Five Exclusive Segments Apply these rules in order: Champions first, Loyal Customers second, Promising Customers third, At-Risk Customers fourth, Hibernating Customers fifth. The order prevents overlaps and provides a fallback for every possible score. Five destinations, no overlaps, nobody left wandering. In this sequence, Champions have R, F, and M of at least 4. Loyal Customers have R and F of at least 3 after Champions are removed. Promising Customers have R of at least 3 and F of 1 or 2. At-Risk Customers have R of 1 or 2 and F of at least 3. Hibernating Customers have R and F of 1 or 2. If R, F, and M are in E2:G2, use =IF(AND(E2>=4,F2>=4,G2>=4),"Champions",IF(AND(E2>=3,F2>=3),"Loyal Customers",IF(AND(E2>=3,F2<=2),"Promising Customers",IF(AND(E2<=2,F2>=3),"At-Risk Customers","Hibernating Customers")))) . Test boundary rows before rollout. A 555 is Champion. A 431 is Loyal because frequency is 3. A 112 is Hibernating. A 353 is Loyal, not Champion. A 235 is At-Risk because inactivity overrides historical spend. Mapping trap: assuming a segment deserves a campaign because it exists. Calculate reachable customers and expected outcomes first. If baseline conversion is 8%, a cell of 500 customers produces about 40 expected conversions; a 10% holdout contains 50 customers and only four expected conversions. That holdout is too thin for confident lift measurement. Pool several monthly cycles, enlarge the holdout, or combine treatments. Assign One Job, Then Measure Migration Champions: protect 90-day repeat rate. Prioritize recognition, service recovery, and benefit use over blanket discounts. Loyal Customers: shorten median time between orders. Compare the new interval with their own prior 90-day baseline. Promising Customers: drive order two within 30–45 days, adjusted to the category’s observed reorder interval. At-Risk Customers: trigger treatment after 1.5 times the customer or category median reorder interval, not an arbitrary calendar date. Hibernating Customers: limit acquisition cost. Use low-cost channels; suppress chronic non-openers. Treatment trap: sending every segment a renamed 15% discount. That changes copy, not strategy. Champions need reliability. Promising buyers need confidence. At-Risk buyers need a timely reason to return. Use 5–10% holdouts only when expected conversion counts support a useful comparison; otherwise rotate treatment by month. Save one snapshot every 30 days. Use columns for customer ID, snapshot date, previous segment, current segment, previous RFM score, current RFM score, reachable status, treatment, holdout flag, next-90-day orders, and next-90-day net revenue. Report upward, flat, and downward migration alongside customer counts. Judge the model over rolling 90-day windows. Clicks and redemptions diagnose campaign execution; they do not prove retention. A Promising customer moving to Loyal is success. An At-Risk customer making one discounted purchase but remaining At-Risk is weaker than the campaign dashboard claims. Keep RFM simple until migration stops explaining revenue differences. Then add one field—category, margin, or predicted reorder date—and test whether it changes action. Ground that decision in the retention math behind LTV, churn, and repeat rate , not demand for a prettier dashboard. --- # Domino’s Loyalty Strategy: A Habit-First Operating Thesis https://loyalflow.cc/blog/dominos-loyalty-strategy-habit-first The short version: Domino’s loyalty strategy supports a habit-first thesis: ordering convenience does the retention work; points reinforce it. Copy that sequence, then prove the economics with second-order rates, contribution margin, and a holdout. Key takeaways Treat Domino’s as an operating model, not proof that an app or points program causes retention. Baseline second-order rate, checkout completion, and identified-order share before changing rewards. Price rewards as a percentage of qualifying spend, then subtract that cost from contribution margin. Judge digital migration over 60–90 days against a baseline or holdout. Fix one repeated ordering task before building another loyalty feature. Domino’s loyalty strategy is a sequence, not a feature list Domino’s spent years making digital ordering useful through saved customer details, remembered baskets, order tracking, and repeat-order shortcuts. Its loyalty proposition sits on top of that ordering infrastructure. That sequence supports the thesis; it does not prove that points caused every improvement in frequency or digital sales. The distinction matters. Public company results combine pricing, promotions, store operations, delivery performance, advertising, menu changes, and channel migration. An operator cannot isolate loyalty impact by pointing at total digital revenue after launch. The attribution trap: calling every app order incremental. A customer who moves from phone ordering to the app may create no new revenue. The migration still has value if it lowers handling cost, improves identification, increases basket size, or produces more repeat purchases. Establish a pre-launch baseline for four measures: checkout completion, identified-order share, second-order rate, and contribution margin per order. Then compare the next 60–90 days with the prior period, a phased rollout, or an unexposed customer group. Without that comparison, the case study is branding, not analysis. The useful Domino’s lesson is therefore narrower and stronger: make the repeated transaction easier before paying for repetition . That claim can be tested in any business without copying Domino’s app, promotions, or program branding. Convenience should improve behavior before points enter the model A loyalty program cannot repair a purchase path customers dislike. Stored payment, saved addresses, clear fees, remembered orders, accurate status updates, and fast reordering create value on every transaction. A reward creates value only when the customer earns or redeems it. Untie the repeat purchase before rewarding it. This week, choose one high-volume task and measure its current completion rate and median time. Good candidates include account sign-in, address entry, payment, basket reconstruction, reward discovery, or order-status lookup. Remove one field, one screen, or one repeated decision before adding another campaign. Use a 30–45-day second-order window when that period covers at least one plausible repurchase cycle for the category. If customers normally buy every 90 days, use 90–120 days instead. The rule is simple: the window must include a realistic next purchase without becoming so long that product changes and promotions obscure the result. The classic failure: reporting downloads and registrations as retention. Those figures show distribution and account creation. Habit evidence appears in completed reorders, shorter reorder intervals, greater use of saved baskets, and fewer abandoned checkouts. Promotional messaging also needs operational discipline. Suppress routine offers while a delivery failure, refund, charge dispute, or unresolved complaint remains open. One badly timed coupon can tell the customer that internal systems do not share context; Lifecycle Suppression Rules: Stop Marketing Through Service Failures gives the practical control logic. Price rewards from margin, not competitor earn rates Do not adopt a 3%, 5%, or 8% reward value because another brand uses it. Calculate the cost directly: reward cost divided by qualifying spend equals reward rate . Then subtract expected reward cost, discounts, payment fees, fulfillment cost, and service recovery from order contribution. The reward budget lives inside the margin. Take a $30 qualifying basket. A 3% reward rate creates $0.90 of face value; 5% creates $1.50; 8% creates $2.40. If the order produces $6 of contribution before loyalty, those rates consume 15%, 25%, and 40% of that contribution respectively, before breakage or incremental behavior is considered. Set a margin floor before launch. Example: if finance requires at least $4.50 contribution from that $30 basket, a fully redeemed $1.50 reward reaches the floor before any extra discount or service credit. An 8% rate breaches it. The correct rate is the richest one that remains above the floor and produces enough incremental frequency to cover redeemed cost. Reward path matters too, but there is no universal purchase count. Model the customer’s normal frequency. A benefit requiring six purchases may feel close for a weekly buyer and irrelevant for a quarterly buyer. Show progress clearly, keep redemption to one or two actions, then test whether members reach the first benefit within one or two normal purchase cycles. The economics failure: funding points from revenue instead of contribution margin. Revenue can rise while profit falls because existing customers receive rewards on purchases they already intended to make. Use a holdout or phased rollout to estimate incremental orders rather than treating every redeemed reward as successful retention. Measure digital migration separately from incremental demand Digital ordering can improve economics without creating a single additional order. Identified transactions support cohort analysis and targeted messaging. Saved preferences can reduce checkout effort. Structured orders may reduce phone handling and manual re-entry. Visible add-ons can affect average order value. A new channel is not automatically a new customer. Measure those effects separately. For identified-order share, calculate identified digital orders divided by total eligible orders. For repeat rate, calculate customers placing another order inside the chosen window divided by first-time customers in the starting cohort. For contribution, use net revenue minus product cost, variable labor, payment fees, discounts, rewards, refunds, and other variable fulfillment costs. Run the comparison for 60–90 days. Report changes in identified-order share, handling cost, checkout conversion, average order value, repeat rate, service failures, and contribution margin. Label channel shifts as migration unless frequency, basket size, cost, or retention improves relative to the baseline or holdout. The measurement trap: combining migration savings and incremental revenue into one success number. They are different value sources with different confidence levels. Report each separately so management can see whether the program created demand, reduced cost, or merely changed where customers ordered. Behavior should outrank stated enthusiasm. NPS vs Repeat Rate: Behavior Proves Retention explains why completed purchases deserve more weight than survey intent. Once point balances become material, add issuance, expiration, redemption, fraud, and liability controls described in Loyalty Points Liability: Build Controls Before Campaigns . Frequently asked questions Does this Domino’s loyalty strategy require an app? No. Use the lowest-friction owned channel customers will revisit. Mobile web, stored browser checkout, an existing commerce account, or a wallet pass may solve the repeated task without native-app maintenance. When should an operator add points? Add points after checkout performs reliably and baseline repeat behavior is known. Launch with a margin floor, a defined reward rate, and a holdout or phased rollout; stop or revise the offer if contribution declines without measurable frequency lift. What should the first weekly audit include? Review checkout completion, payment failures, identified-order share, second-order rate, reward cost, contribution margin, unresolved service cases receiving promotions, and the most common abandonment step. Fix the largest repeated defect before adding features. --- # B2B Loyalty Programs: Design for Renewal, Not Enrollment https://loyalflow.cc/blog/b2b-loyalty-programs-renewal The short version: B2B loyalty programs should make renewal economically obvious before procurement turns the contract into a price comparison. Use account-level rebates for profitable incremental behavior, then give buying teams service and access benefits that support continued adoption. Key takeaways Start with the renewal date, then work backward 90–120 days. Pay rebates on incremental or strategically valuable behavior, not automatic baseline spend. Send economic value to the contracting account; give individual users approved service benefits. Show accrued value, retained status, and next-period economics before competitive bidding begins. Set the maximum reward from contribution economics , not a universal rebate benchmark. B2B loyalty programs win before renewal Consumer loyalty programs optimize repeat transactions. B2B retention has a different decisive moment: contract renewal, rebid, or budget approval. The account may buy regularly, but the real question arrives when procurement asks whether switching suppliers is worth the disruption. Set the renewal route before procurement reaches the junction. Work backward from that decision. For a contract expiring on 31 December, begin the renewal sequence around 90 days earlier. Confirm qualification, quantify earned value, resolve billing disputes, and surface unused benefits before procurement has already defined the conversation around price. The customer dashboard should answer four questions immediately: What did we earn? What did we use? What status do we retain? What disappears if we leave? Keep the last question factual. The objective is not a punitive exit fee; it is a clear comparison between continuing value and switching cost. The classic failure: launching enrollment in January, collecting activity data all year, then mentioning the program during the final renewal call. By that point, the buying committee may have issued a request for proposal. Treat renewal visibility as a lifecycle requirement, not a sales presentation. This week, list every active contract, expiry date, decision group, and current benefit balance. Add a 120-day alert for complex accounts and a 90-day alert for simpler renewals. Use rebates to fund incremental behavior Most B2B buyers already understand invoice credits, volume rebates, marketing funds, and service allowances. Points add a second currency to a relationship already governed by negotiated prices and contract terms. Use points only when their operational benefit clearly exceeds their accounting and redemption burden; otherwise, use a transparent rebate. Reward the new course, not the wall already standing. A rebate rate of 1–5% of qualifying spend can be a useful starting range, not a default promise. Qualifying spend might mean volume above a baseline, adoption of a higher-margin category, multi-year commitment, forecast accuracy, or payment performance. The rule must reward behavior that improves contribution or retention economics. Example: an account normally spends $500,000. It reaches $600,000 after adopting a product family with a 30% gross margin. The additional $100,000 creates $30,000 of gross profit. A 3% rebate on that incremental volume costs $3,000, leaving $27,000 before servicing costs. A retroactive 3% rebate on the full $600,000 costs $18,000 and may pay for demand the account would have generated anyway. Use incremental bands where possible. A rebate on spend from $500,001 to $600,000 protects the baseline. If a cliff is commercially necessary, model the full retroactive cost before approval and cap the exposure. The classic failure: calling an existing discount a loyalty reward. The account receives money but changes nothing. Establish the baseline from trailing 12-month spend, then document why each qualifying action deserves additional value. For low-frequency buying, points create even more friction: progress becomes invisible, redemption takes too long, and the buyer forgets the program between orders. A rebate or contract credit fits the purchasing rhythm better. See loyalty programs for infrequent purchases for the broader case against points in sparse purchase cycles. Separate account economics from buyer enablement The contracting company should receive the economic reward. Apply it as an invoice credit, renewal credit, approved marketing fund, service allowance, or payment to the legal entity. Record the qualifying activity, calculation, approval, and settlement date. Value feeds the account; support shelters the people using it. Individual users still influence adoption and renewal. Give them benefits that improve their work: priority support, training, certification, implementation reviews, advisory sessions, early product briefings, or relevant peer events. These benefits reinforce usage without creating an undisclosed personal payment for purchasing influence. Maintain two views. The account view shows qualified spend, estimated rebate, contract status, service usage, and renewal value. The user view shows training, permissions, support access, and recognition. Do not make a procurement user personally responsible for tracking corporate funds. The classic failure: sending gift cards or expensive personal rewards to employees who control supplier selection. Employer policies, procurement controls, and anti-bribery rules may prohibit them. Route material value to the account. Require documented employer approval for any individual benefit with meaningful cash value. Simple test: disclose the benefit to the buyer's finance director. If the arrangement becomes difficult to explain, replace it with training, access, or account-level value. Build the renewal value statement A rebate balance alone rarely secures a complex renewal. Combine financial and operational evidence: earned credits, products adopted, service consumption, completed training, support outcomes, implementation milestones, and the next period's expected economics. Show the statement at least 60–120 days before expiry . For a 12-month contract, schedule a qualification review around day 245–275, a value review around day 275–305, and commercial negotiation afterward. Timing varies by procurement cycle; the principle does not. Include a plain comparison. “Renewal preserves $X in earned credit, Y active integrations, and Z agreed service capacity.” Do not claim savings without a defensible baseline. Do not hold an earned rebate hostage to signature unless the contract explicitly defined that condition before the account qualified. Retained status may require renewal, committed volume, or an annual business review. Give a 30–60 day grace period for documented contracting delays outside the customer's control. Communicate any status change before the renewal window, never after it. The classic failure: making benefits technically available but operationally invisible. Unused training, unclaimed service, and uncommunicated credits do not create perceived value. Assign an owner for each benefit and report usage before renewal. Set reward limits from contribution economics Do not use a universal rule such as “rewards must stay below 10–20% of incremental gross profit.” The correct ceiling depends on servicing cost, retention value, strategic fit, and the contribution required by the business. The reward comes from the margin, not the whole pie. Use this formula for each account or segment: maximum rebate = incremental gross profit − servicing cost − required contribution . If incremental gross profit is $30,000, servicing costs are $4,000, and the business requires $20,000 of contribution, the maximum rebate is $6,000. A $3,000 rebate works. A $10,000 rebate does not, even if the sales team expects renewal pressure. Review the calculation monthly. Settle rebates quarterly for frequent purchasing; settle annually for seasonal or contract-based volume. Display estimated accrual monthly, resolve disputes within 30 days, and prevent unapproved manual overrides. Measure renewal rate, incremental gross profit after reward cost, share of wallet, product breadth, service usage, and tier or status movement. Enrollment and points issued are activity metrics, not proof of retention. Use behavioral measures alongside satisfaction measures, as discussed in NPS vs repeat rate . The classic failure: treating a large renewal as proof that the program worked. Compare participating accounts with their own trailing 12-month baseline where possible. If revenue rises while contribution falls, the program is subsidizing demand, not retaining profit. Keep the first version narrow: one account-level rebate, one renewal dashboard, one buyer-benefit policy, and one approval formula. Add tiers only when customer behavior, contract complexity, and economics justify them. For model selection across points, tiers, and cashback, use the loyalty model comparison . Frequently asked questions Should B2B rebates be paid quarterly or annually? Use quarterly settlement when purchases are frequent and visible progress supports retention. Use annual settlement for seasonal volume or contract-based buying. Show estimated accrual monthly in both cases. Should rewards go to the account or the buyer? Send economic value to the contracting account. Give individual users approved training, access, recognition, and service benefits. Avoid personal cash equivalents unless the employer has explicitly approved them. How early should renewal value appear? Show accrued value 60–120 days before expiry. Use the longer window for multiple stakeholders, formal procurement, or competitive bids. Start later only when the buying cycle is demonstrably shorter. How do B2B loyalty programs avoid margin loss? Calculate the maximum rebate from incremental gross profit, servicing cost, and required contribution. Pay on incremental or strategically valuable behavior. Audit retroactive cliffs before launch. --- # Lifecycle Suppression Rules: Stop Marketing Through Service Failures https://loyalflow.cc/blog/lifecycle-suppression-rules-service-failures The short version: Lifecycle suppression rules should stop promotional messages when a customer has an unresolved delivery, payment, return, or support problem. Build the pause-and-restart logic before adding another onboarding, cross-sell, referral, or loyalty campaign. Key takeaways Suppress promotions within minutes of a service failure, not during the next daily audience refresh. Use explicit event rules for delays, failed payments, returns, low ratings, and open support cases. Keep transactional updates running while pausing discounts, referrals, reviews, and cross-sells. Restart messaging only after resolution plus a 24–72-hour cooling period. Measure prevented conflicts, post-resolution conversion, complaints, and margin—not email volume. Lifecycle suppression rules need an event hierarchy Most lifecycle systems decide who qualifies for a message. Better systems also decide who must not receive it. A customer with an open delivery problem should be excluded even when that customer qualifies for five revenue campaigns. One unresolved problem outranks every campaign qualification. Start with four event groups: fulfillment failures, payment failures, returns or refunds, and support problems. Useful triggers include a shipment delayed beyond the promised date, one failed delivery attempt, a payment failure, an initiated return, a rating of 1–2 out of 5, or a support case still open after 24 hours. Assign severity before adding channel logic. A missing order or disputed charge should suppress every promotional channel. A minor product question may pause cross-sell for 24 hours while leaving educational usage messages active. The counterexample: a customer reports a missing $120 order at 10:00, then receives a referral request at 14:00 because the campaign audience was built overnight. Both automations worked as configured. The operating rule failed. Use one precedence rule: unresolved service events outrank promotional eligibility. Do not reproduce that decision separately across email, SMS, push, and loyalty tools. Send a single suppression state downstream wherever the current stack permits it. Rule to apply this week: identify the five highest-severity customer events, map every promotional channel they should pause, then test whether the suppression arrives within 15 minutes. If the stack cannot move that quickly, use the shortest reliable sync interval and document the exposure window. Pause promotions, not necessary communication Suppression should not make the company disappear. Customers still need order updates, refund confirmations, security notices, password resets, support replies, and legally required messages. The rule separates necessary communication from messages asking for more money or effort. Pause discounts, product recommendations, replenishment prompts, referral requests, review requests, tier celebrations, points-expiry pressure, and subscription upgrades. Continue messages that explain status, required action, expected resolution, or completed remediation. Classify every automated message as transactional, service, educational, or promotional. This takes 60–90 minutes for a modest program with 20–40 active messages. Anything without an owner or classification should default to promotional until reviewed. The classic failure: a blanket suppression blocks the refund confirmation along with the cross-sell. The customer then contacts support again because the system withheld the one message needed to reduce uncertainty. Transactional labels cannot become a loophole. An email containing a shipping update plus a large “buy again” module is partly promotional. Remove the merchandising block during an active suppression state rather than pretending the whole message is operational. Keep suppression reasons visible to support agents. A simple status such as “promotion paused: return open until resolution” helps agents explain what will happen and prevents manual campaign enrollment during the dispute. Restart after resolution, not case closure A closed ticket does not prove restored confidence. The case may have been closed automatically, the refund may still be pending, or the replacement may not have arrived. Restart conditions should use the customer outcome, not the support team’s administrative status. Fixed is not the same as ready to ring again. For a delivery failure, wait until confirmed delivery or refund. For a return, wait until refund issuance or exchange shipment. For a failed payment, restart after successful payment unless the customer cancelled. For a low rating, require a response or a defined cooling period rather than assuming silence means recovery. Add a 24–72-hour delay after resolution before promotional messages resume. Use the shorter end for simple payment corrections; use 48–72 hours after missing deliveries, damaged products, or disputed charges. Apply a frequency cap so queued campaigns do not all release together. The counterexample: a replacement order arrives on Friday, then three paused campaigns send within 20 minutes: review request, replenishment offer, and points-expiry warning. Suppression prevented the initial conflict but the restart logic created another one. Discard stale messages instead of queueing them indefinitely. A delivery education email may remain useful for 7–14 days. A flash sale ending tomorrow does not. Every paused message needs an expiry condition, even if that condition is simply “skip this send.” Make the first post-resolution message low pressure. Confirm the remedy, provide relevant usage help, or ask whether the issue is actually solved. Do not use a coupon as the default apology; compensation should match failure severity and expected contribution margin. Measure conflicts prevented and value recovered Campaign engagement cannot tell you whether suppression works. Track the number of promotional sends blocked during active service events, the percentage released incorrectly, repeat contacts within 7 days, unsubscribe and complaint rates, and purchase behavior during the 30–60 days after resolution. Audit a sample of 25–50 suppressed customers each week during rollout. Check whether the trigger arrived on time, the correct channels paused, necessary messages continued, and restart happened only after the defined outcome. This manual review catches mapping errors faster than aggregate reporting. Use a holdout only when customer treatment remains fair. Compare normal restart timing with a longer promotional pause; never withhold shipment, refund, or support communication. Judge the result on incremental contribution margin after discounts, returns, and service cost. The classic failure: the team celebrates 10,000 suppressed emails without checking whether those customers later recovered. High suppression volume may indicate good controls, poor fulfillment, or both. The count diagnoses exposure; it does not prove retention. Wait until the final measured cohort has completed the full observation window. A 60-day post-resolution outcome requires 60 days after the last customer enters that cohort, plus any material return window. Leading indicators such as complaints can be read earlier; mature repeat-purchase results cannot. Set an operational target before launch: fewer than 1% of audited customers should receive a prohibited promotion during an active high-severity event. Then test whether post-resolution repeat rate holds or improves without excessive compensation. Suppression protects the measurement behind lifecycle work. Use the retention math behind LTV, churn, and repeat rate to evaluate recovered behavior rather than message activity. Frequently asked questions Should open support tickets suppress every campaign? No. Suppress by severity and subject. A missing order, disputed payment, return, or unresolved product defect should pause promotions. A simple usage question may only pause product recommendations for 24 hours. How long should suppression last? Keep it active until the customer outcome is complete, then add a 24–72-hour cooling period. Set a separate escalation for cases still unresolved after 3–7 days; never restart merely because a timer expired. What if the marketing platform cannot process real-time events? Use the fastest reliable audience sync, then remove the highest-risk campaigns from delayed channels. A 15-minute sync is reasonable for many programs; a 24-hour batch leaves too much room for conflicting messages. Should points-expiry messages continue during suppression? Usually not. Pause expiry pressure during high-severity failures, then extend the deadline by 7–30 days when the customer could not reasonably use the benefit. Keep the adjustment controlled because loyalty points liability requires clear issuance and expiry rules . --- # Loyalty Points Liability: Build Controls Before Campaigns https://loyalflow.cc/blog/loyalty-points-liability-controls The short version: Loyalty points liability requires measurement when a program creates a probable future obligation, but the accounting timing depends on award type, contract terms, and jurisdiction. Build the reconciliation, approval gates, and audit trail before issuing points at scale. Key takeaways Separate purchase-linked awards, promotional grants, and manual credits before applying accounting treatment. Reconcile point movement to the loyalty platform and general ledger every month. Estimate redemption from observed cohorts, not borrowed industry benchmarks. Approve campaigns using expected cost, a high-redemption case, and contribution margin . Keep every assumption, override, and approval in a dated audit trail. When loyalty points liability requires measurement A purchase-linked award can create a performance obligation or deferred-revenue component when the customer earns an enforceable right to a future benefit. A promotional grant issued without a purchase may instead be treated as a marketing expense or provision, depending on its terms. Service-recovery credits, partner-funded points, and discretionary adjustments can require different treatment again. Do not force every point into one accounting bucket. Start with an award-type register containing the earn trigger, funding party, customer right, expiry rule, redemption options, and approved accounting treatment. Finance should confirm the treatment with its accounting advisers; marketing should not infer it from the point label. Operational exposure begins earlier than formal recognition in some structures. Once customers can redeem an award, the program needs to forecast fulfillment and cash demand even if the ledger treatment differs. That distinction prevents an accounting debate from delaying basic cost control. Use expected fulfillment cost for operating forecasts, not customer-facing face value. If 100 million outstanding points have a modeled 70% redemption rate and weighted fulfillment cost of $0.006 per redeemed point, expected cost is $420,000. Those inputs are illustrative; replace them with observed redemption and contracted reward costs. Control failure: every point receives the same value because the platform exports one balance. The ledger then mixes purchase obligations, campaign expense, partner funding, and discretionary credits that should have been tracked separately. Build a monthly roll-forward finance can audit The monthly worksheet needs these fields: award type, opening points, issued points, redeemed points, expired points, manual adjustments, closing points, expected redemption rate, cost per redeemed point, expected cost, ledger balance, and reconciliation difference. Keep the data at award-type level; add cohort month when redemption behavior differs materially. One slipped tooth should never hide in the close. The point formula is simple: closing points equal opening points plus issued points minus redeemed points minus expired points, plus or minus adjustments. Expected cost equals closing points multiplied by expected redemption rate multiplied by weighted fulfillment cost per redeemed point. Reconciliation difference equals modeled expected cost minus the relevant ledger balance. Set an investigation threshold using materiality, not an arbitrary industry percentage. Use the lower of a fixed financial amount approved by finance or a percentage of modeled expected cost. A smaller program might investigate any difference above $5,000; a larger program may use 1% if that produces a lower threshold under its policy. Both formulas and that threshold are runnable in the points liability calculator : it rolls the balance forward, prices it at expected fulfillment cost and a high-redemption case, and flags the reconciliation difference when it exceeds your threshold. Every adjustment needs a reason code, approver, timestamp, and source reference. Separate fraud reversals, customer-service reinstatements, migration corrections, and expired-point reversals. A net adjustment line without evidence is not a control. Control failure: the platform balance reconciles only in total. A campaign over-issues 8 million points while a migration correction removes the same amount, leaving a clean closing balance and two hidden errors. Reconcile movement by type, not just the endpoint. Estimate redemption from cohorts, not benchmarks Calculate observed redemption by earn cohort: redeemed points from that cohort divided by points originally issued to that cohort, adjusted for reversals. Keep purchase-linked base earn, welcome bonuses, multipliers, service recovery, and partner awards separate until data proves their behavior is similar. Measure the behavior your customers actually leave behind. Do not declare eventual redemption from a three-month-old cohort. Use cohorts old enough to cover the program’s normal redemption cycle, then compare cumulative redemption after consistent windows such as 30, 90, 180, and 365 days. Programs with long purchase cycles may need 18–24 months before older cohorts provide a useful maturity anchor. Apply survival or runoff analysis when large balances remain redeemable beyond the observation window. Otherwise, use a documented tail assumption based on the oldest available cohorts. Refresh the estimate quarterly, or sooner after reward repricing, earn-rate changes, expiry changes, or a redemption variance above the threshold set by finance. Breakage is an output of customer behavior and contractual expiry, not a target marketing can select to make economics pass. Recognize it only under the approved accounting policy. Keep operational forecasts showing both expected redemption and a higher-redemption case. Model failure: finance copies a 70% redemption assumption from last year after marketing doubles the earn rate and adds cash-equivalent rewards. The historical cohort no longer represents the current proposition. Segment the new awards and reforecast. Put campaign exposure behind an approval gate Every material promotion should show baseline issuance, incremental issuance, expected redemption, weighted fulfillment cost, high-case cost, expected contribution margin, and funding owner. Use a high case based on internal forecast error or comparable campaigns; without history, test redemption 5–10 percentage points above the working estimate and issuance 10–20% above plan as explicit scenarios, not claimed benchmarks. Price the blast before pulling the trigger. Set gates against your economics. A practical starting policy requires finance approval when a campaign could raise monthly issuance by more than 10%, increase modeled expected cost by more than 5%, or introduce a new reward-cost structure. Tighten those limits when margins are thin or reward funding requires cash settlement. Approval must preserve the submitted assumptions, data extract date, model version, approver, and maximum authorized issuance. During launch, compare actual issuance with the approved cap daily for concentrated events and weekly for longer campaigns. Pause mechanics automatically where the platform supports a hard cap. Judge the campaign against contribution margin, not revenue. The same discipline used in retention economics and repeat-rate decisions applies here: incremental gross profit must exceed reward cost, campaign expense, and likely displacement. Approval failure: creative launches before finance receives the point multiplier rules. Finance can document the exposure afterward, but cannot control it. No approved model, no campaign. Control expiry without manufacturing breakage Expiry limits indefinite exposure only when the customer’s claim legally ends under clear program terms and the accounting policy permits recognition. A new expiry announcement does not erase existing obligations immediately. Retroactive changes require legal review and usually create avoidable service costs. Expiry should end opportunity, not quietly obstruct it. Use a clear rule, commonly 12–24 months from issuance or qualifying account activity when that matches the purchase cycle. Send a reminder at least 30 days before expiry; consider another 7–14 days before expiry for material balances. Show the exact points, date, and available redemption path. Track expiry notices, delivery status, expired amounts, reinstatements, complaints, and manual overrides. A spike in reinstatements means the nominal expiry total overstates the lasting reduction and creates extra support expense. Finance needs the net outcome, not the batch-job output. Policy failure: engineering expires dormant balances without a signed rule, notice evidence, or reinstatement procedure. The apparent reduction becomes complaints, reversals, and an audit problem. For long buying cycles, skip points for infrequent purchases rather than relying on punitive expiry. Frequently asked questions Who should own the loyalty points liability model? Finance should own the model, accounting policy, materiality limits, and ledger reconciliation. Marketing owns campaign forecasts and mechanics; operations or engineering owns platform extracts and movement reconciliation. Named owners should sign each monthly close. How often should assumptions be updated? Reconcile point movement monthly and review redemption, breakage, timing, and reward cost quarterly. Reforecast immediately after material program changes or when actual results breach the approved variance threshold. Do unredeemed points equal profit? No. Unredeemed points can remain customer claims until redemption, valid expiry, or another contractual extinguishment. Expected breakage may affect measurement under the applicable policy, but an outstanding balance is not automatically profit. What evidence should an auditor receive? Provide program terms, award classifications, monthly roll-forwards, platform-to-ledger reconciliations, cohort calculations, reward-cost support, adjustment logs, campaign approvals, expiry evidence, assumption changes, and dated sign-offs. --- # Loyalty Programs for Infrequent Purchases: Skip Points https://loyalflow.cc/blog/loyalty-programs-infrequent-purchases The short version: Loyalty programs for infrequent purchases should skip points. When purchases sit 3–10 years apart, useful service preserves permission, referrals create interim value, and durable records help the brand win when replacement intent returns. Key takeaways Reject points when meaningful redemption takes longer than customers will remember the account. Build contact around ownership events, not campaign quotas. Cap referral rewards using allowable acquisition cost, not a generic percentage. Track reachable records by purchase cohort and improve the baseline each quarter. Measure service, referrals, recognition, and eventual category repurchase. Why loyalty programs for infrequent purchases fail with points Points require repeated transactions. A customer buying coffee weekly can see progress after several visits. A customer replacing a mattress, boiler, vehicle, roof, or premium appliance every 3–10 years cannot. Points that cannot arrive in time are not much of a reward. Take a $2,000 purchase with a proposed 2% earn rate. The account receives $40 in value, then sees no natural earning event for years. That balance creates neither habit nor switching cost. It is a delayed discount attached to an account the customer may forget. Use a cycle-relative test instead of an arbitrary redemption deadline. Estimate time to first meaningful redemption from actual purchase frequency and spend. Reject points when that time exceeds either the normal repurchase interval or the period during which customers still recognize and use the account. The decision rule is simple. Use points for frequent, measurable purchases. Use service benefits when ownership creates recurring needs. Use access when availability, priority, or expertise has standalone value. Use referrals when satisfied owners can generate demand before buying again. The broader choices appear in Points, Tiers, or Cashback: Choosing the Right Loyalty Program Model . The classic failure: points expire after 12 or 24 months while the category repurchase cycle lasts five years. Customers cannot earn enough to redeem, then discover that the small balance vanished. The program adds liability, support work, and irritation without changing behavior. Build the program around the ownership cycle The program needs a useful job after checkout. Map installation, registration, setup, warranty, maintenance, inspections, replacement parts, seasonal preparation, repairs, resale, and disposal. Keep only events where contact can prevent cost, save time, or improve product performance. Loyalty lasts longer when the product does. Start with a test cadence, not an industry benchmark. For example: onboarding within 7 days, a setup check after 30 days, a maintenance reminder at the product’s documented interval, and a warranty notice 60 days before expiry. Measure action rate and unsubscribes, then remove messages that produce neither service activity nor retained permission. Give every contact one action. Book service. Retrieve the correct manual. Confirm coverage. Order a compatible part. Store an inspection record. A message without a clear ownership outcome does not deserve space in the lifecycle. Make the customer record useful enough to revisit. Store model, serial number, purchase date, warranty status, invoices, parts, and completed work. Test retrieval internally: choose a target such as 10 seconds, measure current performance, then fix the largest identity or search failure. The classic failure: replacing purchase frequency with marketing frequency. Monthly newsletters and generic tips preserve a sending schedule, not a relationship. If two ownership events merit contact this year, send two useful messages rather than 24 forgettable ones. Use referral economics, then preserve customer identity A referral can act as the repeat transaction between major purchases, but it does not need a second loyalty program. Define one qualifying event: an attributable introduction that becomes a paid, non-cancelled customer. Ignore shares, clicks, leads, and quotes unless their downstream economics are proven. Let advocacy travel; keep the customer record rooted. Set the reward from allowable customer acquisition cost. Use this formula: maximum referral reward = allowable CAC − handling cost − expected fraud cost − expected cancellation cost . If allowable CAC is $300, administration costs $20, expected fraud costs $15, and cancellation exposure is $25, the reward ceiling is $240. Test below that ceiling; do not default to a percentage of order value. Choose one attribution method and a test window, such as a named referral lasting 60 days. Pay only after the cancellation period. Review repeated claims manually once observed fraud or support cost justifies the work. This keeps the referral layer small and tied to profitable demand. Meanwhile, preserve identity across email changes, phone changes, dealers, installers, and service partners. Match consented identifiers with product serial number, order number, address, or warranty registration. Do not force duplicate accounts because the original transaction came through a channel partner. Define reachability as deliverable, permissioned customer records ÷ eligible purchase cohort . Establish the baseline by purchase year and channel. Set the next quarterly target as a measured improvement, such as 3 percentage points, rather than claiming that one universal threshold fits every category. The classic failure: paying referral rewards for low-quality leads while customer records decay. Self-referrals consume budget, sales teams dispute attribution, and obsolete contact details make the eventual replacement campaign useless. One conversion event, one reward rule, one identity owner. Measure the relationship across the full replacement cycle Monthly active members mean little in a category bought once per decade. Measure whether the program keeps customers reachable, produces useful ownership behavior, creates profitable referrals, and improves recognition when category demand returns. Measure the whole relationship, not one quiet season. Track service uptake after each reminder, warranty registration, successful record retrieval, reachable-record rate, referral conversion, referral contribution margin, assisted revenue, and category repurchase. Keep cohorts based on original purchase year, product type, and channel for the full expected cycle. Compare customers who used service or completed a referral with similar customers who did neither. The comparison does not prove causation, but consistent differences in reachability, consideration, and repurchase tell operators where to test next. Enrollment alone proves only that checkout staff asked or an incentive worked. The classic failure: reporting a 40% enrollment rate as retention while ignoring dead email addresses, unused benefits, and absent repurchase data. Behavior remains the harder standard, as explained in NPS vs Repeat Rate: Behavior Proves Retention . This week, kill the points proposal, map five ownership events, define one referral conversion, calculate its reward ceiling, audit reachable records by cohort, and create a quarterly dashboard. The program should become quieter, easier to operate, and more useful. Frequently asked questions How often should an infrequent-purchase brand contact customers? Contact them when an ownership event creates a useful action. Start with onboarding, maintenance, warranty, safety, parts, or inspection events. Measure action and unsubscribe rates; remove contacts that produce neither customer value nor service activity. How large should a referral reward be? Start below the economic ceiling: allowable CAC minus handling, fraud, and expected cancellation costs. Use a fixed amount when clarity matters. Pay after the cancellation period and tighten review only when claim volume or observed abuse warrants it. How do you measure loyalty with a 10-year replacement cycle? Use leading indicators while repurchase matures: reachable records, service uptake, warranty activity, referral contribution margin, assisted revenue, and recognition during category research. Preserve purchase cohorts for the full cycle rather than resetting reporting annually. What if customers buy through dealers or marketplaces? Offer a direct ownership benefit worth registering for, such as warranty coverage, service history, parts lookup, or maintenance reminders. Record the originating channel, preserve dealer credit, and assign responsibility for consent, service contact, and replacement follow-up. --- # NPS vs Repeat Rate: Behavior Proves Retention https://loyalflow.cc/blog/nps-vs-repeat-rate The short version: In the NPS vs repeat rate debate, repeat rate wins. Purchases prove retention; NPS supplies hypotheses about why customer behavior changed. Key takeaways Define repeat rate using a fixed 30-, 60-, 90-, or 180-day window matched to the purchase cycle. Compare mature acquisition cohorts before reviewing NPS movements. Keep survey trigger, channel, wording, and timing stable. Link responses to later purchases, then control for tenure, channel, and previous frequency. Use behavioral confirmation for retention investment , not urgent safety, fraud, accessibility, or compliance fixes. NPS vs repeat rate: behavior wins NPS asks whether a customer would recommend the company. Repeat rate records whether that customer returned and bought again. One captures stated intent at a specific moment; the other captures an economically meaningful action. Retention leaves wear marks, not promises. Choose the repeat window from the natural purchase cycle. Monthly consumables may need 30- and 60-day views. Apparel may need 90 or 180 days. Compare customers acquired in the same period, then give every cohort equal observation time. A customer who scores the brand a 10 but never returns produces no retained revenue. A customer who scores it a 6 but buys four times in six months does. When sentiment and transactions disagree, transactions decide whether retention happened. The classic failure: celebrating an NPS increase from 42 to 47 while 90-day repeat rate falls from 28% to 22%. The survey may have reached happier customers, followed successful support cases, or missed silent defectors. None of those explanations repairs the six-point decline. Use one primary window and one secondary window. Review 30- and 90-day repeat for faster categories; 90- and 180-day repeat for slower ones. Fix those definitions for at least two reporting cycles rather than changing the window when results become uncomfortable. Survey design can manufacture an NPS trend NPS respondents are not a random customer sample. Customers with unusually good or bad experiences often respond more readily. Quiet defectors may disappear from survey data precisely when their behavior matters most. The survey may catch the loudest customers, not the whole market. Timing changes the measure. A survey sent minutes after delivery mainly captures delivery satisfaction. One sent after a refund captures service recovery . A survey sent 30–45 days later may better reflect product use, but usually attracts fewer responses. Channel changes create another distortion. Email, in-app, receipt, and support surveys reach different populations. Comparing an in-app score this month with an email score last month mixes customer selection with sentiment. Sample size also limits interpretation. With 100 responses, a simple proportion near 50% has an approximate 95% sampling margin of error near ±10 percentage points under random sampling. NPS combines promoter and detractor shares, while voluntary response bias adds uncertainty that this calculation cannot remove. The classic failure: treating 20 angry responses as proof that the whole customer base is leaving. Those comments can expose serious friction, but they do not establish prevalence. Check returns, support contacts, purchase delays, and repeat behavior before funding a broad retention intervention. Keep the survey channel, trigger, wording, and delay stable for 8–12 weeks when possible. That period is an operating heuristic, not a statistical guarantee: it usually provides enough cycles to separate persistent movement from a single campaign, outage, or fulfillment issue. Always show response counts beside the score. Measure the behavior behind the score Start with cohort repeat rate: the percentage of first-time buyers who place another order inside the chosen window. Add median time to second purchase, orders per returning customer, and retained revenue from the original cohort. Each measure answers a different question. Repeat rate shows how many customers return. Time to second purchase shows how quickly a habit forms. Purchase frequency measures depth among returners, while revenue retention catches customers who remain active but spend less. Keep the scorecard compact. For each monthly acquisition cohort, record the relevant 30-, 60-, 90-, or 180-day repeat rates. Add median days to order two, orders per returning customer, and retained revenue as a percentage of first-order cohort revenue. Segment only where an operator can act: acquisition source, first product, geography, membership status, or first-order value band. Five usable segments beat 40 cuts with tiny denominators. For formulas and cohort setup, use the retention math every founder should know . The classic failure: reporting one blended repeat rate after acquisition mix changes. If paid social grows from 20% to 50% of new customers, total repeat rate can fall even when every channel remains stable. Compare like-for-like cohorts before blaming the product or loyalty program. Do not declare a 90-day retention shift from a cohort aged 45 days. Require full observation time, then look for either two consecutive mature cohorts moving in the same direction or one large movement confirmed across meaningful segments. Two cohorts are a practical guardrail against reacting to one noisy period, not proof of causation. Use NPS to diagnose, not certify NPS becomes useful when linked to behavior. Compare scores and comments among fast repeaters, late repeaters, one-time buyers, high-value customers, and customers who returned products. Ask which experience changed before behavior moved, not whether the headline score rose. NPS detects a signal; purchases confirm the diagnosis. Review mature cohort repeat rate, time to second purchase, frequency, and retained revenue first. Flag a movement for investigation when it persists across two periods or is large enough to affect the operating plan. Avoid a universal 10% threshold; normal volatility differs sharply between a cohort of 200 customers and one of 20,000. Trigger surveys around specific experiences: 3–7 days after delivery, after support resolution, or after enough usage time to judge the product. Cap requests near one every 60–90 days to reduce fatigue and prevent frequent buyers from dominating responses. Store customer ID, order ID, trigger, response date, score, and consent status. Append purchases occurring 30, 60, and 90 days later. Compare later behavior across score bands while controlling for tenure, acquisition channel, and prior purchase frequency. The classic failure: seeing detractors repeat less and claiming NPS caused churn. Poor experiences can drive both outcomes; product fit, delivery region, or customer type may also explain the relationship. NPS identifies where to investigate. It does not establish causation. Require behavioral confirmation before committing substantial retention budget. Do not wait for repeat-rate evidence to address credible safety, fraud, accessibility, or compliance problems; those demand immediate investigation and containment regardless of revenue impact. The same behavior-first test applies to events and brand experiences: experiential loyalty needs a second habit, not a packed room . Positive comments matter only when they lead to another valuable action. Frequently asked questions How many NPS responses are enough? Calculate precision from the decision being made. Around 100 responses gives roughly ±10 percentage points for a proportion near 50% under random sampling; voluntary survey bias makes real uncertainty worse. Combine periods for small segments, inspect comments for hypotheses, and always report the denominator. How often should NPS and repeat rate be reviewed? Review survey themes and available behavior weekly. Make retention decisions monthly or quarterly, matched to the purchase cycle. A 90-day repeat metric cannot support a trustworthy weekly verdict. What if NPS rises while repeat rate falls? Trust the repeat-rate warning. Check cohort maturity, acquisition mix, survey channel, response rate, respondent composition, and purchase-window definitions. Treat higher NPS as a diagnostic clue, not proof that loyalty improved. How should responses be linked to purchases? Attach each response to a stable customer ID and timestamp under appropriate consent, access, and retention controls. Compare purchases before and after the response, then segment by tenure and prior frequency to avoid mistaking established loyalty for survey impact. --- # Experiential Loyalty Needs a Second Habit, Not a Packed Room https://loyalflow.cc/blog/experiential-loyalty-second-habit Experiential loyalty fails when the experience ends at the door. Tinder Events may fill rooms, generate social posts, and lift app opens for several days. None of that proves retention unless attendees return to another event, continue useful conversations in the app, or remain paying customers longer than comparable non-attendees. The standard is repeat behavior. If Tinder cannot connect first attendance to another valuable action within 30–45 days, Events is an acquisition campaign—not a loyalty program. Experiential Loyalty Starts With a Recurring Job An event must solve a customer problem that returns. For Tinder, that problem cannot be broad social discovery. It needs a recognizable recurrence window: meet compatible singles nearby, restart dating after a quiet week, enter a trusted social setting, or move stalled conversations offline. A memorable night becomes loyalty only when it leads somewhere again. The job determines the cadence. A neighborhood mixer every two or four weeks can become routine. A celebrity launch party cannot. A recurring interest-based meetup can build familiarity; an annual festival mainly creates a memory. The classic failure: treating one attendance as retention. A customer who attends once, posts twice, then stops using Tinder has not become more loyal. The event bought temporary attention. Test the job before expanding the format. Interview 10–15 attendees within 72 hours. Ask what problem the event solved, when that problem is likely to return, and what would make another booking worthwhile without a discount. Do not impose an arbitrary repeat-intent threshold. Start with historical behavior: how often eligible users already attend singles events, how often Tinder can offer a relevant local event, and how many repeat attendees are needed to cover fixed delivery costs. Then write a hypothesis before launch, such as: attendees offered one relevant event within 30 days will repeat more often than a comparable holdout group receiving no event invitation. Small samples need ranges, not declarations. Report the repeat-rate estimate with its confidence interval. If 8 of 40 attendees return, the observed rate is 20%, but the uncertainty remains too wide for a large rollout. Run more cycles before treating the result as stable. Build the Event Activation Funnel Tinder should operate Events as a product funnel: discovery, detail-page view, RSVP, attendance, post-event connection, second-event offer, then repeat attendance. Each stage needs one owner, one definition, and a weekly cohort review. The room matters less than the thread that brings people back. RSVP-to-attendance should be judged against Tinder’s own event history, split by free versus paid entry, lead time, city, venue distance, and reminder cadence. A free RSVP made 21 days ahead is not comparable with a paid booking made three days ahead. Blending them hides the actual leak. Measure post-event connection within 72 hours, while people still remember names and conversations. Define it tightly: a reciprocal match, exchanged messages, or another mutually agreed interaction. Profile views and app opens are too weak; both can rise without creating customer value. Second-event booking needs an eligible-opportunity denominator. Suppose a city runs one event every 30 days, but only 60% of first-time attendees qualify for the next format by age, location, or interest. Repeat rate should be reported for all first-time attendees and separately for those with a relevant second opportunity. Showing only exposed users exaggerates performance; showing only the full cohort can hide poor event supply. The funnel trap: optimizing registration because it is easy to move. A shorter form can lift RSVPs while attendance, connections, and repeat bookings remain flat. Fix the first behavioral leak, not the most visible interface. Use one 45-day cohort view. Compare users first exposed during the same week, then track every stage. If attendance is weak, change commitment devices, timing, or reminders. If connections are weak, change the room design and matching flow. More promotion will not repair a weak event. Measure App Retention and Break-Even Economics The primary behavioral metric should be repeat attendance within the next eligible event window. For a biweekly format, that may mean 30 days. For a monthly format, use 45–60 days. The window must allow at least one realistic chance to return. Events must also strengthen Tinder’s core product. Compare attendees with similar non-attendees by city, tenure, prior activity, and subscription status. Track meaningful behavior during the next 14 days: reciprocal matches, replies, active conversations, profile improvements, or intentional browsing. Raw app opens reward aggressive notifications. Paid effects need 30–60 days of observation. Separate new subscriptions, renewals, and prevented cancellations because each has different economics. A burst of upgrades can look attractive while subscriber churn remains unchanged. Replace generic churn targets with a break-even rule. Calculate total event cost, including venue, staffing, safety, incentives, support, and allocated local operations. Divide that cost by eligible attendees to get the required incremental retained margin per attendee. If an event costs $12,000 and reaches 300 eligible attendees, it needs $40 of incremental retained margin per attendee to break even. That value may come from additional renewals, lower cancellations, or profitable upgrades. Ticket revenue can reduce the cost base, but it should not be mistaken for retention value. Use a randomized holdout where practical. If not, use matched non-attendees and state the limitation. Compare incremental retained margin over the same period, then attach a confidence interval. Scale only when the lower end of the plausible range approaches break-even—not when the point estimate briefly clears it. The measurement failure: presenting registrations, impressions, or social reach as evidence of loyalty. Those metrics describe distribution. They do not establish changed behavior, reduced churn, or profitable retention. Close the Event-to-App Loop, Then Kill Weak Formats The event should create a useful reason to reopen Tinder within 24 hours. With mutual consent and clear privacy controls, the app can surface people met at the venue, conversation prompts, confirmed connections, or the next relevant local event. Keep the next action close. A follow-up should arrive within 24–72 hours; the next event option should appear within 7–14 days when supply permits. Booking should take fewer than three taps. Five screens, an unrelated home feed, or a generic thank-you message breaks continuity. No points system or blanket 20% discount is required. The reward is better access, trusted attendance, improved matching, and continuity between online and offline interaction. Discounts can test price sensitivity, but they cannot prove recurring demand. Give each format an 8–12 week test window, long enough for several occurrences and at least one repeat opportunity. Predeclare the decision rule: continue when incremental retained margin plausibly covers cost; redesign when one funnel stage clearly blocks repeat behavior; stop when three cycles produce no durable app or retention lift. The final trap: protecting a photogenic format because it fills the room. Crowds may be first-time visitors with no intention of returning. Attendance earns another test only when repeat behavior and cohort economics improve. Events deserve funding when the second habit pays for the first experience. Use LoyalFlow’s framework for LTV, churn, and repeat rate to calculate the retained margin Tinder Events must produce before expansion. --- # Burger King’s Whopper Guarantee Is a Retention Bet, Not a Refund Policy https://loyalflow.cc/blog/burger-kings-whopper-guarantee-is-a-retention-bet-not-a-refund-policy A bad Whopper can erase the next 5, 10, or 20 visits from a customer who decides the brand is unreliable. Replacing one burger matters only if it prevents that loss. Burger King should therefore treat the Whopper Guarantee as a retention mechanism, not a refund policy. The operating target is not replacements issued. It is customers recovered. The Guarantee Must Repair the Current Meal Product failures create an immediate decision: was this a random mistake, and will Burger King fix it without a fight? Fast replacement answers both questions before frustration becomes a reason to switch. The replacement has to rescue today’s meal before it can earn tomorrow’s visit. Use five minutes as a pilot target, not an industry benchmark. Start the clock when the customer reports the problem. Stop it when the corrected item reaches them. Measure the current median by location, then test whether frontline authority can reduce it over four weeks. A voucher delivered seven days later may reimburse the food, but it does not repair the meal. Points have the same limitation. Value redeemable on a future visit asks the customer to accept today’s failure and take another risk later. Set one decision rule: visible product-quality failures receive immediate replacement below a defined order-value limit. Managers handle ambiguous claims or repeated requests. Crew members should not need approval for a clearly incorrect, missing, cold, or badly prepared item. Failure signal: replacement time looks acceptable only because refused claims never enter the dataset. Log every request, including denials and abandonments. A low claim rate without denial data proves nothing. This week, choose 5–10 restaurants with different order volumes. Record acknowledgment time, resolution time, outcome, channel, failure type, and whether manager approval was required. That produces an operating baseline without pretending five minutes is universally correct. Use Store Economics, Not Generic Food-Cost Claims A replacement burger does not cost its menu price, but Burger King-specific direct cost cannot be inferred from public menu pricing. Operators should use their own ingredient, packaging, waste, and incremental labor data. The useful equation is simple: replacement direct cost ÷ preserved contribution margin per future visit . If replacement costs $2 and a normal future visit produces $4 of contribution margin, preserving one visit covers two replacements. Those figures are illustrative assumptions, not Burger King estimates. Run the calculation at location level. Menu mix, labor conditions, franchise economics, and delivery fees can change the answer materially. Use the customer’s normal order history where identity matching exists; otherwise use channel-specific average contribution margin. Set review thresholds from the pilot rather than importing arbitrary percentages. Establish each location’s weekly claims per 1,000 eligible orders, then investigate stores running at twice the pilot median or moving sharply for two consecutive weeks. Review patterns before restricting legitimate claims. The classic failure: finance attacks visible replacement cost while acquisition discounts remain buried in marketing spend. Compare both on the same basis: direct cost per retained or acquired customer, followed by contribution margin over 30 and 60 days. Do not require a full lifetime-value model. Start with the next two expected visits. If the recovered customer returns once at normal contribution margin, the guarantee may already pay back; if return behavior remains depressed, a cheap replacement was still a failed recovery. Build One Claim Record, Then Measure the Next Visit The minimum measurement schema needs one row per request. Store claim ID, customer or payment identifier where permitted, order ID, location, channel, failure type, request time, resolution time, outcome, replacement direct cost, manager involvement, and denial reason. Track the repair, then watch whether the customer comes back. Join that record to four behavioral fields: pre-incident visit frequency, pre-incident average spend, first return date, and spend during the next 30 and 60 days. Preserve an anonymous cohort for customers who cannot be matched rather than excluding their operational results. Build the baseline from the 60–90 days before launch. For each claimant, estimate expected visits using their own prior frequency when available. A customer who normally visits weekly should not be judged by the same 60-day standard as someone who visits quarterly. Use a compact weekly view: Operations: claims per 1,000 orders, median resolution time, first-contact resolution, denial rate, manager involvement. Quality: failure type, repeat complaint rate, location variance, recurring menu-item problems. Retention: 30-day and 60-day return rate, days to next visit, post-incident spend versus prior spend. Economics: replacement direct cost, preserved contribution margin, cost per recovered customer. A useful pilot target is relative improvement. Reduce median resolution time or denial rate by 25% over four weeks, then check whether 30-day return behavior improves. Relative targets remain defensible because they use Burger King’s own baseline. Counterexample: 90% of claims receive replacements, yet recovered customers return half as often as before. Operational completion looks strong; retention remains damaged. Compare post-incident frequency with both matched non-claimants and each customer’s prior behavior. Review execution weekly by location, customer cohorts monthly, economics quarterly. System averages hide the store where every claim requires a manager and the store where one equipment problem generates repeated failures. Keep Recovery Separate From Loyalty Rewards Points reward continued behavior. Guarantees repair broken behavior. Combining them lets loyalty mechanics obstruct a basic service obligation. A customer holding 800 points does not need another 200 when the burger is wrong. Replace the product first. Add points only as extra recognition when the failure involved unusual delay, repeated mistakes, or another reason the replacement alone was insufficient. Membership should never determine eligibility. Digital identity can simplify order matching and later retention analysis, but forcing enrollment turns recovery into lead capture. Honor the guarantee first; invite enrollment afterward. Keep claim intake under 60 seconds as a design hypothesis. Known digital orders should require only failure type, requested remedy, and optional evidence. Counter claims can use approximate purchase time, item, and location when the defect is visible. Failure signal: customers must upload multiple photos, verify an email, or wait 48 hours for a low-value decision. That process may reduce claims while increasing churn. Fraud controls should target repeated or anomalous behavior, not tax every claimant. Pilot the workflow across 5–10 varied locations for four weeks. Expand only after Burger King can identify who approves replacements, how quickly stores resolve them, what each replacement costs, and whether resolved customers return at something close to their prior frequency. Guarantees earn budget through preserved contribution margin, not free-food volume. Model that trade with the retention math every founder should know , then judge the Whopper Guarantee by the visit that proves recovery worked: the next one. --- # Win-Back Emails That Actually Win: A Lifecycle Playbook https://loyalflow.cc/blog/win-back-emails-that-actually-win-a-lifecycle-playbook By the time most brands send a win-back email, the customer stopped being a customer months ago. The playbook below is built on one principle: win-back begins before churn, not after . Define "lapsed" from your own data Forget generic 90-day rules. Pull the distribution of gaps between orders; your risk threshold is roughly the 80th percentile of that gap. If 80% of repeat purchases happen within 45 days, a customer at day 50 is already unusual — that is when the sequence starts, not at day 90 when the habit is gone. The four-touch sequence The nudge (at risk threshold). No discount. Best-sellers, what's new, a reason to visit. You are testing whether attention, not price, was the problem — and protecting margin on everyone this recovers. The reason-why (7–10 days later). Address the actual objection: restock reminders for consumables, social proof for considered purchases, sizing/fit help for apparel. Segment if you can; even two variants beat one generic blast. The offer (10–14 days later). Now the discount — single-use, expiring , and meaningful (10% rarely moves a lapsed buyer; 15–20% with a deadline does). One offer, one deadline, no stacking. The goodbye (30+ days later). "We'll stop emailing." Honest, and it reliably produces a last spike of recoveries while cleaning your list — deliverability is a retention asset too. Measure incrementality or measure nothing Hold out 10% of each lapsed segment from the entire sequence. Revenue per recipient versus holdout is the only number that justifies the discounts. Typical result: the nudge and goodbye emails outperform expectations, and the discount step is doing less work than it appears — which is exactly why you test. Win-back is the highest-leverage flow in lifecycle marketing because the audience already trusted you once. Treat it as a system with a clock, not a coupon with a subject line. Setting the clock from your own data Pull every customer with at least two orders and list the gap in days between consecutive orders. Sort those gaps and read the 80th percentile. If 80% of repeat orders land within 38 days, 38 is your risk threshold — not 90, and not whatever the last agency deck said. Now hang the sequence off it. The nudge goes at day 38, the reason-why around day 46, the offer at day 58, the goodbye at day 90. Every date derives from one measured number, which is what makes the schedule defensible when someone asks why the discount fires when it does. Two things distort that percentile, and both matter. Customers with a single order are not in the distribution at all, so the number describes people who already repeat: it is a threshold for at-risk repeaters, and one-time buyers need a different flow entirely. And seasonal categories produce a gap distribution with two humps rather than one — if you see that, split the calculation by season instead of averaging into a threshold that fits neither. Recompute quarterly until the number stops moving. A threshold nobody has checked in a year is a generic rule with extra steps. Frequently asked questions I do not have enough order history to compute the 80th percentile. Now what? Use the distribution you have and mark the threshold as provisional. With a few hundred repeat orders you can still see where the mass of gaps sits, even while the exact percentile keeps moving. What does not work is importing a 90-day rule from another category: the gap distribution for coffee and the one for mattresses have nothing to say to each other. Should the goodbye email actually unsubscribe them? It should stop the marketing stream, because that is the promise it makes. Keep transactional and service messages running — different basis, different expectation. Suppressing rather than deleting also preserves the record you need if that customer returns through another channel and someone has to explain why they stopped hearing from you. Do consumables and considered purchases use the same sequence? Same shape, different clock and a different second touch. For consumables the risk threshold is a replenishment interval and the reason-why message is a restock reminder, which often recovers the customer before any discount is needed. For considered purchases the normal gap is long enough that lapse and patience look identical, so the sequence matters less than fixing why the second purchase had no trigger in the first place. --- # What Starbucks Rewards Gets Right (and What Most Copycats Miss) https://loyalflow.cc/blog/what-starbucks-rewards-gets-right-and-what-most-copycats-miss Starbucks Rewards routinely drives more than half of U.S. company-operated revenue, and its stored-value balances rival a small bank's deposits. Every retailer has tried to copy it. Most copy the surface — stars, an app, free drinks — and miss the machine underneath. 1. The prepaid float is the program The genius is not points; it is stored value . Members load money onto cards before buying anything. That produces three effects copycats rarely replicate: Starbucks holds billions in interest-free float, breakage on unspent balances flows to revenue, and — most important behaviorally — money already loaded feels spent, so the next purchase decision is pre-made in Starbucks' favor. 2. Rewards priced in perceived value, not cost A free latte costs Starbucks well under a dollar in marginal ingredients but is priced to the member at $5–6 of value. That gap lets the program feel generous at a modest true cost. Retailers who sell other people's products at thin margins cannot reproduce this — which is why a supermarket copying the stars model ends up either stingy or unprofitable. 3. Frequency mechanics, not annual ones Coffee is a daily habit , and every mechanic matches that cadence: stars expire in months, double-star days create short-term urgency, challenges reset weekly. The lesson is not "add gamification" — it is match reward cadence to purchase cadence . A mattress brand with a punch card has copied the form and ignored the physics. 4. The app is the loyalty program Order-ahead, payment, and rewards live in one surface, so the program is not a discount layer — it is the most convenient way to buy. Convenience is the retention mechanism; the stars are the story members tell themselves. What to steal Prepaid or subscription mechanics if your frequency supports them — the float and pre-commitment do the heavy lifting Rewards with a perceived-value-to-cost gap (your own products, experiences, access) Expiration and cadence tuned to your natural purchase cycle Copy the physics, not the paint. The copy that fails, in order The pattern repeats: a retailer ships stars, an app and a free item, then wonders why frequency did not move. It comes apart in a predictable sequence. The reward is funded from a thin resale margin, so it is set low enough to be uninspiring — a $5 reward after $500 of spend is a 1% rebate wearing a costume. Because the reward is distant, members stop tracking progress. Because nobody is tracking progress, the app has no reason to be opened between purchases. Because it is not opened, it never becomes the way to buy, and it stays a loyalty screen bolted onto a checkout that already worked. Notice what is absent at every step: money in advance. Without pre-commitment, the program can only react to purchases the customer was already making, which is the definition of a discount rather than a retention mechanism. So the diagnostic is one question. Does anything in your program change what the customer does before they decide to buy? Stars awarded afterwards do not. A loaded balance, a paid membership or a subscription does. Frequently asked questions Can a smaller brand run stored value? Mechanically, yes — most commerce platforms already sell gift-card or store-credit functionality. The parts that stop people are not mechanical. Prepaid balances are customer money until spent, which brings gift-card and unclaimed-property rules that vary by jurisdiction, and breakage recognition is an accounting policy rather than a marketing decision. Get both answered before you promote top-ups, not after the balances exist. Is breakage just revenue from customers who forgot? Partly, which is why the recognition policy matters. Breakage taken aggressively converts a customer-service problem into reported revenue, and it tends to come back later as complaints and refunds. Unspent is also not the same as forgotten: a customer who tops up every month permanently carries a float they fully intend to use. Estimate breakage from your own observed redemption cohorts rather than a borrowed rule of thumb. My customers buy monthly, not daily. What still carries over? The pricing logic and the pre-commitment, not the cadence mechanics. A reward with a wide gap between perceived value and marginal cost works at any frequency. Star expiry, weekly challenges and double-point days do not — they assume enough purchase occasions for urgency to land somewhere. Match the mechanic to your actual interval, or you are running a countdown the customer cannot beat. --- # The Retention Math Every Founder Should Know: LTV, Churn, and Repeat Rate https://loyalflow.cc/blog/the-retention-math-every-founder-should-know-ltv-churn-and-repeat-rate Loyalty conversations go wrong when they run on vibes — "engagement," "delight," "community." The businesses that get retention right run on four numbers. None of them are complicated; all of them are routinely computed wrong. 1. Repeat purchase rate (RPR) The share of customers who buy a second time. For most e-commerce, a 20–30% RPR is typical; above 40% is strong. It is the single most honest indicator of product-market fit for retention, because no incentive program can rescue a product nobody wants twice. 2. Churn — measured on a cohort, not a blend Blended churn hides everything. If you acquired heavily last month, your "average" churn looks great while every cohort is quietly leaking. Always read churn as: of customers acquired in month X, how many were still active in month X+n? Plot three cohorts and you will learn more than from a year of blended dashboards. 3. LTV — with margin, not revenue The most common LTV inflation: using revenue instead of contribution margin. A $300 revenue LTV at 25% margin is a $75 customer. If acquisition costs $60, you are running a very tight boat while your dashboard celebrates. Loyalty rewards come out of that margin too — a 2% earn rate on a 25%-margin business consumes 8% of your profit pool. 4. Payback window How long until a cohort's cumulative margin covers its acquisition cost. Under 6 months and you can reinvest aggressively; over 18 and growth is financed on hope. Retention programs move this number more reliably than any acquisition optimization, because every extra order lands inside an already-paid-for relationship. The uncomfortable conclusion A loyalty program is a margin reallocation: you tax every transaction to change future behavior. It pays back only if the incremental orders it creates exceed the discounts it hands to customers who would have returned anyway. Estimating that incrementality — not the point balance, not the signup count — is the entire game, and it is why every serious program needs a holdout group from day one. If you want to run these numbers against your own figures, the retention calculator does the arithmetic above, including what one percentage point of retention is actually worth to you. The four numbers on one business Take a store with 1,000 customers acquired in January, an average order value of $80, and a 25% contribution margin after payment fees, fulfilment and returns. Repeat purchase rate: 240 of them buy a second time, so RPR is 24%. That sits inside the normal band — nothing here is broken, and nothing is exciting either. Cohort churn: by month six, 180 of the original 1,000 are still buying. Read that as a cohort retained at 18% after six months, not as "82% churn", which sounds like a crisis and describes nothing you can act on. LTV: those repeat customers average 3.2 orders at $80. Revenue LTV reads $256. Contribution LTV is $64. If acquisition cost $45, the business works — barely, and only if that margin number is honest. Payback: at $20 of contribution per order, $45 of acquisition cost clears after roughly two and a quarter orders. On a six-week buying interval that lands in month four, which is inside the range where reinvesting is defensible. Now add a loyalty program earning 2% of spend. That is $1.60 per $80 order against $20 of contribution — 8% of the margin pool, charged on every repeat order including the ones that needed no incentive at all. The program has to create enough additional orders to cover that, which is the entire reason the holdout exists. Frequently asked questions How many cohorts do I need before cohort churn means anything? The constraint is cohort size, not cohort count. A cohort of forty customers moves several points of retention when two people leave, so the noise is larger than the signal you are looking for. Plot three consecutive cohorts and read the shape rather than the decimal. If your monthly volume is small enough that single-digit churn swings the line, widen the cohort to a quarter instead of reading noise as a trend. Gross margin or contribution margin in the LTV calculation? Contribution: gross margin minus the costs that scale with an order — payment fees, fulfilment, shipping subsidy, returns, and the reward itself. Gross margin flatters LTV precisely because it hides the costs that grow with every extra order a loyalty program buys you. If contribution margin is not available yet, run the number both ways and treat the gap between them as the size of your uncertainty rather than picking the friendlier one. How large does the holdout need to be? Large enough to detect the effect you would act on, which depends on your base rate and that effect — not on a fixed percentage. Work backwards: if repeat rate is 25% and you would only keep the program for a three-point lift, the holdout has to be big enough for three points to be distinguishable from ordinary variation. If that sample exceeds your monthly volume, run the holdout for longer rather than shrinking it, and accept that you are measuring a quarter rather than a month. --- # Points, Tiers, or Cashback: Choosing the Right Loyalty Program Model https://loyalflow.cc/blog/points-tiers-or-cashback-choosing-the-right-loyalty-program-model Most loyalty programs fail before launch, at the moment someone says "let's just do points." The model you choose is not a branding decision — it is an economic contract with your customers, and each of the three dominant models makes a different promise. Points: flexible, but easy to get wrong Points work when purchase frequency is high and order values vary. Grocery, coffee, beauty — anywhere a customer buys weekly and can "save up" toward something meaningful. The two levers that matter are earn rate (the effective discount, usually 1–5% of spend) and burn friction (how easy redemption feels). The classic failure: an earn rate so conservative that the first reward sits six months away. If a new member cannot see a realistic path to a reward within 30–45 days, the program is dead on arrival — the liability sits on your balance sheet while the motivation never materializes. Tiers: status for high-variance spend Tiers shine when a minority of customers drive a majority of revenue — airlines, hotels, fashion. You are not paying for transactions; you are paying for identity . Silver, Gold, Platinum work because losing status hurts more than earning it feels good. That loss aversion is the engine. Rule of thumb: your top tier should be reachable by roughly the top 5–10% of customers. Any looser and status means nothing; any tighter and nobody plays. Cashback: simple, honest, expensive Cashback is the bluntest instrument: a transparent rebate, usually 1–3%. It converts well because there is nothing to explain, but it buys no emotion and no switching cost — the moment a competitor offers 4%, your "loyalty" evaporates. It fits low-margin, high-competition categories where simplicity is the differentiator. How to decide High frequency, moderate margin (coffee, grocery, pharmacy) → points Concentrated revenue, aspirational brand (travel, fashion, B2B ) → tiers, often layered on points Commodity category, price-driven buyers (fuel, electronics, marketplaces) → cashback The model is the skeleton. The next question — how generous to be — is where the real margin math starts, and that deserves its own article. The same business under all three A specialty grocer: $45 average basket, 22% contribution margin, customers buying roughly twice a month. Points at 2% of spend cost $0.90 a basket against $9.90 of contribution — about 9% of the margin pool. But run the reward path before congratulating yourself: a $15 reward takes seventeen baskets to reach, which at twice a month is over eight months. That breaks the 30-to-45-day rule this article opened with. The fix is a smaller, faster reward — around $3 — or a higher earn rate you have actually funded. Both are decisions; drifting into a distant reward is not. Tiers on the same business mostly do not work. Spend is concentrated around a routine fortnightly shop rather than spread across a long tail, so a top tier reachable by the top 5-10% would separate customers by a handful of baskets a year. Status that easy to reach, and that undramatic to lose, does not produce the loss aversion tiers exist to create. Cashback at 2% costs the identical $0.90 and buys less: an instant rebate with no progress to track, no threshold to reach and nothing forfeited by shopping elsewhere next week. What it does buy is comprehension, which is worth something when the category is price-led and the comparison happens at the shelf. Same cost, three different behaviours purchased. The model question is not which one is generous — it is which mechanic your purchase pattern can support. Frequently asked questions Can I combine two models? Yes, and tiers layered on points is the usual pairing: points do the transactional work, tiers do the identity work. What to avoid is three currencies a customer has to hold in their head at once. If you cannot state the whole program in one sentence at the till, it is too complicated to change behaviour, whatever the feature list says. How do I actually set the earn rate? Start from what you can fund, not from a competitor rate. Take contribution margin per order, decide what share of it you are willing to spend on repeat behaviour, and work back to a percentage of spend. A 2% earn rate in a 25%-margin business commits 8% of the profit pool on every transaction, including those from customers who were returning anyway. Then check the other end: at that rate, how many orders until a reward is reachable? If the answer runs past a couple of months of normal buying, the rate is too low to motivate anyone regardless of what the model says. What if my margins cannot fund any of these? Then the answer is not a thinner version of the same program. Non-discount benefits — priority access, service levels, useful information, faster support — cost operations rather than margin, and a competitor cannot match them by quoting a bigger number. A thin-margin business running 1% cashback has bought the right to be outbid.