The Experimentation Edge

The Experimentation Edge

Growthbook
Pays États-Unis
Langue EN
Épisodes 43
Dernier 22.09.2026

The Experimentation Edge is a podcast about how product teams decide what to build and what not to build. Product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive business outcomes, with examples from companies like DoorDash, Atlassian, and UPS. Hosted by Ashley Stirrup, CMO at GrowthBook, it is aimed at product managers, engineers, data scientists, and growth leaders at B2B tech companies. Episodes focus on experimentation culture, statistical rigor, and shipping with confidence, without marketing speak.

Épisodes

  • Why GoPro stopped judging A/B tests by win rate 22.09.2026 26min
    Will Guyeskey is Director of Digital Product at GoPro, where his team owns the e-commerce side of gopro.com. Before GoPro he ran personalization at Gap and cut his teeth at Brooks Bell, testing for brands like Barnes & Noble, Under Armour and Ralph Lauren. In this episode of The Experimentation Edge, he tells Ashley Stirrup why knowing your customer is the through line of every good test program.Will shares the Barnes & Noble order confirmation test he was sure would lose, and why the same idea never worked for any other client. He walks through a recent GoPro Mission launch test that asked whether a step-by-step configurator adds too much friction, and what a flat result revealed about high consideration buyers. He also explains why GoPro shares interim readouts across the company, why win rate makes a poor North Star for an experimentation program, and how his team plans to use AI for speed without outrunning its own learnings.Chapters00:00 Intro00:52 Will's role running e-commerce at GoPro01:43 Learning A/B testing across retail at Brooks Bell05:24 How GoPro runs one to three tests a month06:35 Sharing learnings and interim readouts across teams09:20 The Barnes & Noble order confirmation win13:02 Why the win did not transfer to other clients14:41 Testing friction on the GoPro Mission configurator18:51 Why win rate is the wrong North Star22:19 How AI will shape experimentation at GoProTakeaways- A winning idea rarely travels. The Barnes & Noble recommendation module worked because of that audience's low order values and reading habits, and it failed for every other client that tried it.- Design every test so it teaches you something whether it wins, loses or ends flat. Losing tests are jet fuel when the learning is built in.- A flat result is still an answer. GoPro's configurator test showed that buyers of high consideration products accept extra steps when each choice adds value.- Share interim readouts across the company, and use them to show how volatile results are before a test reaches statistical significance.- AI can speed up building and running experiments, but a team that runs more tests than it can learn from is not getting better.Connect with the GuestWill Guyeskey LinkedIn: https://www.linkedin.com/in/willguyeskey/About the guest: https://www.growthbook.io/podcast/guests/will-guyeskeyEpisode page: https://www.growthbook.io/podcast/episode/1-43Company Website: https://gopro.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io
  • What Samsung learned bringing B2C rigor to B2B 17.09.2026 30min
    SummaryWhat changes when you take a mature B2C experimentation practice and apply it to a B2B storefront? Anuradha Tempe, Lead Product Manager at Samsung Electronics America, joins host Ashley Stirrup to share what she learned managing both sides of Samsung.com. She explains why B2B is not a sidekick to B2C, how a single bulk order can push a test into a false positive unless you normalize the data, and why not everything deserves an A/B test. She walks through the add-on experiment that nearly doubled attach sales once her team realized business buyers decide on mobile and purchase on desktop, and the Buy Now test that lost because B2B buyers value clarity over faster conversions. The conversation closes with her approach to North Star and guardrail metrics and the AI copilot she built to draft A/B test plans with a human still in the loop. A practical episode for product managers, engineers, data scientists, and growth leaders running experimentation across different customer types.Chapters00:45 Meet Anuradha Tempe: from chip design to leading Samsung e-commerce02:15 Why B2B is not a sidekick to B2C03:30 Not everything needs an A/B test05:15 Spreading experimentation practice across a global conglomerate07:00 Bringing B2C rigor to a fast-paced B2B team08:45 The add-on experiment that nearly doubled attach sales11:15 Buyers decide on mobile and purchase on desktop16:30 The Buy Now test that lost21:30 North Star metrics, guardrails, and the EPP discount fix26:05 An AI copilot for A/B test planningTakeaways- Treat B2B as its own customer base with its own testing discipline; a single bulk order on one day can inflate a B2B test into a false positive unless order data is normalized before results are read.- Not everything needs an A/B test; route lower-risk changes through UAT feedback or pre/post comparisons and reserve full experiments for features where being wrong is expensive.- Map where the decision happens, not just where the purchase happens; Samsung's business buyers decide on mobile and buy on desktop, and surfacing add-ons on mobile nearly doubled attach sales.- B2B buyers value clarity over faster conversions; a Buy Now button earlier in the flow confused bulk purchasers because returns and cancellations on large orders are costly.- Align on the North Star before building, whether it is revenue, engagement, NPS, or fewer support tickets, and set guardrails so an engagement feature can never quietly drag sales down.Connect with the GuestAnuradha Tempe LinkedIn: https://www.linkedin.com/in/anuradha-tempe/About the guest: https://www.growthbook.io/podcast/guests/anuradha-tempeEpisode page: https://www.growthbook.io/podcast/episode/1-42Company Website: https://www.samsung.com/us/business/SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Even a loss is a win: Charlie Health's approach to experiments 15.09.2026 30min
    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Joe Yevoli, Director of Growth at Charlie Health, a virtual intensive outpatient program that sits between weekly therapy and hospitalization. Joe explains how removing a page from Charlie Health's intake form produced a winning test that created a new bottleneck further down the funnel, why Teachers Pay Teachers cut about 80% of its market for one feature and roughly quadrupled retention, and how premortems that ask "what went wrong?" before launch make dissent safe and surface safeguards that avoid catastrophe. He also shares the two questions he asks before any experiment, why more top-of-funnel traffic always dents conversion rate, and how AI is flattening the product pod in ways that speed teams up and put them at risk. It's for growth leaders, product managers, and experimentation teams who want to learn as much from a loss as from a win.Chapters00:00 Intro01:00 About Charlie Health02:05 Joe's role across performance, lifecycle, and experimentation03:30 Building experimentation rigor across teams04:40 Even a loss is a win05:15 The form page test that created a new bottleneck08:50 Teachers Pay Teachers and the Easel lesson16:30 Designing experiments that teach you something when they lose21:00 Premortems for high-risk experiments24:30 AI is flattening the orgTakeaways- A winning test is a data point, not a finish line; Charlie Health's form page removal won on top-of-funnel metrics and still exposed a downstream bottleneck that became the next experiment.- More traffic entering the funnel means conversion rate goes down, even when more people reach the bottom; treat it as a law of physics and plan for it before you read the results.- Build it for everyone and you build it for nobody; Teachers Pay Teachers cut roughly 80% of the market for Easel, repositioned it around one job, and roughly quadrupled retention.- Run a premortem before any large, risky experiment: ask the room "it was a massive failure, what went wrong?", have everyone write and share, and build the safeguards while there is still time.- Before launching, check the quantitative and qualitative data behind the hypothesis, map the full funnel, and measure toward the real business outcome, so a loss still leaves you with a next step.Connect with the GuestJoe Yevoli LinkedIn: https://www.linkedin.com/in/joeyevoli/About the guest: https://www.growthbook.io/podcast/guests/joe-yevoliEpisode page: https://www.growthbook.io/podcast/episode/1-41Company Website: https://www.charliehealth.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Realtor.com on using your AI as a junior data scientist 10.09.2026 20min
    SummaryWhat do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up.Chapters00:00 Cold open and welcome01:10 From growth hacker to Realtor.com02:50 Three foundations for a new experimentation team04:20 The 30/30/30 rule05:05 The 300% bundling win that lost revenue07:10 You don't need a stats degree to experiment09:20 Cascading North Star metrics12:50 Do the homework before the experiment13:50 The wishlist: instrumentation, embedded knowledge, culture15:50 AI as an aspirational data scientist18:45 Keeping a human in the loopTakeaways- A winning decision metric is not enough. Realtor.com's bundling test hit a 300% attach rate, but funnel fallout from the extra step made it a net revenue loser. Set secondary metrics and their thresholds before launch.- Expect the 30/30/30 rule: roughly a third of tests win, a third are inconclusive, and a third lose. The math is the math, and the losers carry most of the learning.- Start a new team on foundations: what a clean test and an A/A test look like, which surfaces should not be tested, and a peer review program that lets people graduate to more complex experiments.- You don't need a stats background to run good experiments. Teach the simplest definition of a good test, then let people learn by doing.- AI can make anyone an aspirational data scientist for analyzing results and spotting opportunities, but it can be confidently wrong. Keep a human in the loop and sanity check output the way you'd peek at a freshly launched test.Connect with the GuestWhitney Perez LinkedIn: https://www.linkedin.com/in/whitneykperez/About the guest: https://www.growthbook.io/podcast/guests/whitney-perezEpisode page: https://www.growthbook.io/podcast/episode/1-40Company Website: https://www.realtor.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Zalando connects every experiment to its North Star 09.09.2026 20min
    SummaryHow do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition.Chapters00:45 About Zalando and its marketplace model02:00 Mi's path from engineering to experimentation03:10 Economists and data scientists in one decision-making org04:15 Running over 1,000 experiments a year05:55 What makes an experiment high risk07:10 Sharing learnings through standardization and champions09:10 The discovery feeds launch and its measurement plan11:15 Balancing the experimentation portfolio12:35 Growing a KPI tree from the North Star16:05 LLM agents and the future of experimentation at Zalando18:35 New missions for a longstanding businessTakeaways- Treat experimentation as a portfolio: balance confirmatory tests that protect the business with game-changing bets that can win big.- Assess risk tiers when building the roadmap so measurement rigor scales with the stakes instead of slowing every decision down.- When a launch is too new to have a success metric, pair short term A/B tests with long term holdouts from day one.- Connect every experiment to the company North Star by cascading it down to sensitive proxy metrics and controllable inputs.- Use LLM based agents as a cheap simulation layer to screen hypotheses, not as a replacement for real A/B tests.Connect with the GuestMi Tian LinkedIn: https://www.linkedin.com/in/mi-tian-941b4767/About the guest: https://www.growthbook.io/podcast/guests/mi-tianEpisode page: https://www.growthbook.io/podcast/episode/1-39Company Website: https://www.zalando.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Learneo on testing the opposite of every hypothesis 08.09.2026 21min
    SummaryRich Liebling, senior director of engineering at Learneo, joins host Ashley Stirrup to explain the practice that came out of growing Shop It To Me from 50,000 subscribers to one million in nine months: test your hypothesis, and test its opposite. Rich covers the page where cutting text lost and adding text won, why the inverse wins more often than teams expect, and why small focused tests are the only ones where "the opposite" means anything. He also walks through translating a million subscriber goal into a target of 300 A/B tests, making a new engineer's second merge request their own experiment, and the different constraints at Course Hero, where competing team metrics were resolved with an early lifetime value model and three-day SQL analyses quietly capped testing velocity. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program.Chapters00:00 Cold open and introduction01:50 Shop It To Me and the first engineering hire02:55 300 A/B tests as the path to one million subscribers04:00 Testing as a core value and the second merge request07:20 Learneo, Course Hero, and two ways to get access08:50 Modeling lifetime value to stop teams competing11:10 Testing the opposite of the hypothesis14:35 A portfolio of small, medium, and large tests17:15 Multi armed bandits and seasonal traffic20:35 The flywheel that keeps copycats behindTakeaways- Test the opposite of every hypothesis. At Shop It To Me the inverse won surprisingly often, and even when it lost it proved the variable mattered.- Keep tests small and focused. Redesign a whole page and lose, and you learn that version failed but not what to change next.- Translate a growth goal into an execution count. One million subscribers is not actionable. 300 A/B tests by year end is, and everyone can influence it.- Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea.- Analysis friction sets the ceiling on testing velocity. Three days of ad hoc SQL per test quietly discourages teams from running more.Connect with the GuestRich Liebling LinkedIn: https://www.linkedin.com/in/richliebling/About the guest: https://www.growthbook.io/podcast/guests/rich-lieblingEpisode page: https://www.growthbook.io/podcast/episode/1-38Company Website: https://www.learneo.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • The four questions Early Warning asks before any A/B test 03.09.2026 21min
    SummaryWhat separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs.Chapters00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments02:00 Wayfair and optimizing every step of the storefront funnel03:05 The four questions to ask before any A/B test05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration07:40 Why 85 to 90% of tests fail and why that's a learning agenda09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem12:55 Pre-registration, kill criteria, and stopping p-hacking15:10 Rolling out winners: gradual ramps and long-running holdouts17:15 Causal inference when you can't A/B test20:05 The case for more A/B testing, not lessTakeaways- Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis.- Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy.- Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways.- Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win.- Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs.Connect with the GuestPriya Singhee LinkedIn: https://www.linkedin.com/in/priya-singhee/About the guest: https://www.growthbook.io/podcast/guests/priya-singheeEpisode page: https://www.growthbook.io/podcast/episode/1-37Company Website: https://www.earlywarning.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Supercell A/B tests 300 million players without breaking trust 01.09.2026 23min
    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Shan Huang, data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, Brawl Stars, Hay Day, and Boom Beach. Shan explains how a famously decentralized, creative-first company with 300 million monthly active users runs fewer than 100 A/B tests a quarter and why that restraint is deliberate, how Supercell announces experiments to players in advance and promises make-up events to keep testing fair, and how importable AI skills now let anyone at the company analyze their own experiments, making quality consistency the next big challenge. It's for product managers, data scientists, and growth leaders balancing creative conviction with experimental rigor.Chapters00:00 Intro01:10 About Supercell and 300 million players02:30 The central team and a decentralized culture04:25 Fewer than 100 tests a quarter05:40 Sharing learnings across independent game teams07:35 Retention as the North Star08:35 Onboarding experiments with gems and tutorials12:45 Telling players about A/B tests15:15 Hypotheses and proxy metrics20:45 AI and the future of experiment analysisTakeaways- Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume.- In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams.- Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner.- Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged.- AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company.Connect with the GuestShan Huang LinkedIn: https://www.linkedin.com/in/cnshanhuang/About the guest: https://www.growthbook.io/podcast/guests/shan-huangEpisode page: https://www.growthbook.io/podcast/episode/1-36Company Website: https://supercell.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Clover experiments when billions of dollars flow through daily 27.08.2026 30min
    SummaryHow do you run an experimentation program when classic A/B testing is off the table? Ben Schein, Director of Product Management at Clover, joins host Ashley Stirrup to explain how a platform serving 300,000+ merchants and processing billions of dollars daily proves every feature through pilots and ground-level testing before rollout, why uncertainty and downside — not feature visibility — decide testing depth, and what his years leading product at Shake Shack taught him about turning a checkout funnel into a brand channel. This episode is for product managers, engineers, and data scientists building experimentation programs where the stakes are real.Chapters00:00 Cold open and welcome01:30 The Clover business model and its scale03:45 Ben's role and the metrics that matter07:45 Deciding what gets tested: uncertainty and downside09:05 Why Clover can't test in production12:45 Testing the Shake Shack checkout experience19:35 Advice for PMs new to experimentation23:45 Context over personalization25:45 The future of experimentation and the human elementTakeaways- At Clover's scale, testing happens before rollout: pilots and detailed go-to-market plans replace in-production A/B tests, because a merchant's work tool can never change overnight without warning.- Uncertainty and downside set the testing depth. High-risk changes like payment authorization flows earn deep, ground-level experimentation, while table stakes features like Apple Pay earn a monitored rollout.- Testing is the evidence that justifies rollout investment: if the data doesn't show a feature will succeed, the go-to-market dollars never get spent.- A checkout funnel can carry the brand. At Shake Shack, Ben's team tested prep-time expectations, fixed wrong-location orders, and used loading screens to deliver hospitality digitally.- Structure experiments for durable business value, not pass-fail verdicts: pair headline metrics with counter metrics and anchor on ground-level measurements like items per check that resist marketing noise.Connect with the GuestBen Schein LinkedIn: https://www.linkedin.com/in/benschein/About the guest: https://www.growthbook.io/podcast/guests/ben-scheinEpisode page: https://www.growthbook.io/podcast/episode/1-35Company Website: https://www.clover.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Why JobLeads says one test won't move you, but 100 will 26.08.2026 26min
    SummaryWhat happens when the most logical feature you've ever built has zero impact? In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Edd Saunders, product experimentation manager at JobLeads, to unpack the pizza personalization experiment that cut an ordering flow from 22 clicks to 5 and changed nothing. Edd shares the problem mapping framework he uses to move new experimenters from solution space to problem space thinking, how JobLeads grew from 0.3 to 2.8 experiments per month, and why democratizing experimentation across a whole company comes down to habit change rather than education. A practical conversation for product managers, data scientists, engineers, and growth leaders building experimentation cultures.Chapters00:00 Introduction01:40 Meet Edd Saunders and JobLeads03:12 How JobLeads uses AI for prototyping04:05 Building experimentation operations and a knowledge base05:35 Learning over winning and compounding growth08:10 The pizza personalization experiment12:50 Moving from solution space to problem space17:05 Problem mapping on a 2x2 matrix22:20 Velocity and democratizing experimentation24:20 AI automation for the unsexy workTakeaways- A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control.- Exploration is part of the customer's delight; returning customers wanted to browse the menu even though they ordered the same thing every week.- Moving new experimenters from solution space to problem space thinking raises win rates and produces learnings the whole organization can use.- Problem mapping on a 2x2 matrix of evidence versus impact turns customer research into a prioritized experiment roadmap, and one validated problem can spring a whole tree of testable ideas.- Scaling experimentation from 0.3 to 2.8 tests per month is less about education and more about habit change, shared learnings, and giving non specialists the tools to launch their own experiments.Connect with the GuestEdd Saunders LinkedIn: https://www.linkedin.com/in/eddsaunders/About the guest: https://www.growthbook.io/podcast/guests/edd-saundersEpisode page: https://www.growthbook.io/podcast/episode/1-34Company Website: https://www.jobleads.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Why The Aspen Group targets a 25% win rate 20.08.2026 23min
    SummaryWhat does a winning A/B test mean when your website serves 1,100 dentist owned offices? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arie Polycarpou, Senior Manager of Test and Learn at Aspen Dental, about building experimentation programs at Kohl's, Marriott, Total Wine, and now the largest company in The Aspen Group's healthcare retail portfolio. Arie explains why appointment bookings are the North Star but never the whole story, why he deliberately targets a 25% win rate and expects it to fall as the program matures, and why he applies a haircut to every stacked win before it reaches leadership. A practical conversation for product managers, engineers, data scientists, and growth leaders building test and learn programs at multi-location businesses.Chapters00:00 introduction and Arie's path into experimentation02:05 building programs at Kohl's, Marriott, and Total Wine03:15 healthcare retail and 1,100 dentist owned offices04:30 the test and learn team and 100 tests a year07:15 building a culture of shared wins and explained losses09:50 what a losing navigation redesign revealed12:30 testing your way into big redesigns13:55 the case for a 25% win rate15:50 haircuts, holdouts, and honest math on stacked wins17:55 metrics beyond conversion and where AI fits nextTakeaways- Aspen Dental runs experimentation as healthcare retail: with 1,100 dentist owned offices, the office, not just the website visitor, is the real unit of analysis.- A mature program should target a true win rate around 25%; a 40% win rate usually signals a young program still picking off low-hanging fruit.- Apply a haircut to stacked wins: lifts depreciate as customers acclimate, and two 5% wins never add up to 10%.- Losing tests are valuable when they're designed to isolate why: test your way into big redesigns instead of shipping them whole.- Experimentation culture grows from sharing wins and explaining losses: keep the statistical rigor in the back end and the communication simple.Connect with the GuestArie Polycarpou LinkedIn: https://www.linkedin.com/in/ariepolycarpou/About the guest: https://www.growthbook.io/podcast/guests/arie-polycarpouEpisode page: https://www.growthbook.io/podcast/episode/1-33Company Website: https://www.aspendental.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Principal A/B tests on customers who don't exist 18.08.2026 31min
    SummaryWhat if you could A/B test on customers who don't exist before spending a single live impression? Erika Dunn, assistant director of data science at Principal Financial Group, joins host Ashley Stirrup, CMO at GrowthBook, to share how she built synthetic digital audiences entirely in house: profiles shaped by real data that rank content by likelihood of engagement, whose first live A/B test selection just beat the control. They also dig into the looping metric her team built in SQL to find where customers get stuck without heat-mapping tools, why testing gets watered down into "let's try something" at large companies, and how a center of excellence that shares wins and losses defeats the "we tried that years ago" reflex. This episode is for product managers, data scientists, marketers, and experimentation leaders, especially those working inside large, risk-averse organizations.Chapters00:00 Cold open and welcome01:40 Erika's path from quantitative psychology to experimentation03:35 When testing gets watered down06:50 Experimentation in Principal's marketing space09:55 The looping metric that finds stuck users12:45 Building synthetic digital audiences15:15 The first synthetic audience A/B test wins22:30 The PDF lesson and meeting customers on mobile25:15 North star metrics and the right contact cadence28:05 Where experimentation goes next with AI agentsTakeaways- Synthetic digital audiences let teams rank 20 content options by predicted engagement before spending a single live impression, and Principal's first synthetic selection beat the control in a real A/B test.- A looping metric built from web behavior data can find where customers get stuck without heat-mapping tools: watch how often users cycle back to the same page within tight time windows.- Testing gets watered down when "let's try something" replaces a control group; a little pre-planning gets far more out of every experiment.- A center of excellence that shares wins and losses turns tribal knowledge into shared knowledge and stops "we tried that years ago" from killing valuable retests.- An experimentation mindset requires that people can't get punished for mistakes; give teams guardrails and a safe playground and they'll stop running the same test forever.Connect with the GuestErika Dunn LinkedIn: https://www.linkedin.com/in/erikadunn/About the guest: https://www.growthbook.io/podcast/guests/erika-dunnEpisode page: https://www.growthbook.io/podcast/episode/1-32Company Website: https://www.principal.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Why US Bank considers missing even 1% of customers unacceptable 11.08.2026 22min
    SummaryHow does a major bank scale experimentation when even one percent of customers missing an experience is unacceptable? Vijay Lal, Lead Product Manager for Experimentation at US Bank, joins host Ashley Stirrup, CMO at GrowthBook, to share how his team made their experimentation platform self serve for non technical marketers, how a login widget experiment led to a two second fallback that accounted for every customer, and why metrics should be driven by hypotheses instead of handed down by leadership. They also dig into where AI genuinely saves time in experiment analysis, why a human in the loop is non negotiable, and what real time personalization means for the future of testing. This episode is for product managers, data scientists, and experimentation leaders, especially those working in regulated industries.Chapters00:00 Cold open and welcome00:40 Vijay's path from Comcast to financial services03:16 Making the experimentation platform self serve05:08 The login widget experiment and the two second fallback08:59 Documenting learnings from every experiment10:38 AI in experimentation and the human in the loop12:39 Advice for new product managers15:24 Hypothesis driven metrics18:25 Real time personalization and agentic AI20:11 Democratizing experimentation with responsibilityTakeaways- Self serve experimentation lets a small central team support a huge testing volume, but it only works with continuous training and guardrail metrics attached.- In a regulated industry, every customer must be accounted for. Even one to two percent of users missing an experience is unacceptable.- A simple fallback, like a two second load rule, can save an ambitious experiment without sacrificing coverage or security.- Metrics should be driven by the experiment's hypothesis, not chosen by leadership in a silo. Pair a primary KPI with secondary KPIs for return behavior.- AI saves real time in experiment analysis, but a human in the loop must validate anything AI produces before it goes live.Connect with the GuestVijay Lal LinkedIn: https://www.linkedin.com/in/vijay-lal/About the guest: https://www.growthbook.io/podcast/guests/vijay-lalEpisode page: https://www.growthbook.io/podcast/episode/1-31Company Website: https://www.usbank.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Why Farfetch manages by learning rate, not win rate 05.08.2026 40min
    SummaryLuis Trindade, Principal Product Manager of Experimentation at Farfetch, joins host Ashley Stirrup to explain how one of the world's largest luxury marketplaces built its own experimentation platform and the culture around it. Luis covers the move from a hybrid setup with an external testing vendor to Fabs 2.0, the in house system where a feature toggle is the single entry point for every experiment, why Farfetch manages by learning rate instead of win rate, and the two year Inspire experiment that replaced the world's leading recommendation engine vendor. He also shares how a deliberately shrinking center of excellence supports hundreds of experiments a month through clinics, shared templates, and open learning sessions. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program.Chapters00:00 Cold open and introduction01:45 Inside Farfetch, the global marketplace for luxury fashion08:00 From startup validation to an experimentation mindset09:45 A center of excellence that enables instead of executes12:45 Fabs, build versus buy, and dropping the external vendor16:45 One feature toggle as the entry point for every experiment20:45 Learning rate over win rate23:15 The two year experiment that replaced the recommendation vendor29:45 Onboarding new product managers into experimentation33:15 AI, corporate knowledge, and what comes next for experimentationTakeaways- Manage by learning rate, not win rate. The only failed test is one that was badly designed, with wrong metrics or sampling biases. Every other test produces a learning.- Route every experiment through a single entry point. Farfetch's feature toggling system connects segmentation, user systems, CMS, and messaging so every team tests in the same language.- External JavaScript injection tools carry hidden costs: broken pages, inconsistent results, and rework to reclaim your own data for deep dives.- Strategic bets deserve a longer clock than fail fast allows. Farfetch iterated on its Inspire engine for two years before it beat and replaced the market leader.- A center of excellence should enable, not execute. Farfetch's central team shrank while experiment volume grew because its job is ceremonies, templates, and coaching.Connect with the GuestLuis Trindade LinkedIn: https://www.linkedin.com/in/ltrindade/About the guest: https://www.growthbook.io/podcast/guests/luis-trindadeEpisode page: https://www.growthbook.io/podcast/episode/1-30Company Website: https://www.farfetch.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Cogniteer Built an Experimentation Engine From Scratch 23.07.2026 33min
    SummaryIn this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume.Chapters00:00 Introduction01:25 From agency mass production to in house deep dives04:05 Why some products resist selling online06:35 The drop offs every ecommerce shop shares07:45 The 50% win rate test Cogniteer reused09:25 Why alignment beats developer resources12:45 Two teams, two goals, one broken checkout17:05 Selling water dispensers without a product catalog22:15 Designing every experiment to lose26:15 Personalizing buyers and where AI takes experimentationTakeaways- Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients.- Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product.- Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide.- The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation.- Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog.Connect with the GuestFabian Hans LinkedIn: https://www.linkedin.com/in/fabianhans-cogniteer/About the guest: https://www.growthbook.io/podcast/guests/fabian-hansEpisode page: https://www.growthbook.io/podcast/episode/1-29Company Website: https://www.cogniteer.de/SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Fin A/B Tests Millions of Samples in Days 21.07.2026 45min
    SummaryOn this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.Chapters00:00 Welcome and what Fin actually does02:00 How Fin became Anthropic's first line of support02:30 Why Fin sells resolutions not deflections06:00 Owning the stack with custom models10:40 Pedro's path from fuzzy logic to AI12:55 Why A/B testing is the only gold standard for AI15:50 Do no harm testing on every change18:00 The latency experiment that shocked the team27:30 When more context made Fin hallucinate30:15 Win rates and the future of AI driven experimentationTakeaways- Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work.- You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped.- Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain.- A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.- Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program.Connect with the GuestPedro Tabacof LinkedIn: https://www.linkedin.com/in/tabacof/About the guest: https://www.growthbook.io/podcast/guests/pedro-tabacofEpisode page: https://www.growthbook.io/podcast/episode/1-28Company Website: https://fin.aiSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • How Kargo turns losing experiments into competitive edges 14.07.2026 22min
    SummaryIn this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure.Chapters00:00 Welcome and introducing James Falzone01:45 What Kargo does and how real time ad auctions work04:45 Why experimentation is embedded in Kargo's culture07:45 The three things every marketplace has to deliver10:15 The experiment that failed: click optimization on third party demand12:15 A bad result versus a bad experiment13:45 Why different customer types need different signals15:30 Putting "where did you fail?" on every retro18:45 How experimentation evolves with AI21:15 Better not bigger: the closing takeawayTakeaways- A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new.- The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech.- Metrics and signals you test against should always be business driven, not ported from the last thing that worked.- Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut.- AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger.Connect with the GuestJames Falzone LinkedIn: https://www.linkedin.com/in/jamesafalzone/About the guest: https://www.growthbook.io/podcast/guests/james-falzoneEpisode page: https://www.growthbook.io/podcast/episode/1-27Company Website: https://kargo.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments 09.07.2026 27min
    SummaryIn this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Olean, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data.Chapters00:45 Meet Danielle Olean and Box's reinvention02:45 Owning the entire customer life cycle04:45 Why experimentation matters even without a checkout07:45 The feature that's used but hidden11:45 Proving ROI with a scrappy manual test12:45 Building a culture that shares wins and losses16:45 The pyramid strategy for prioritizing tests18:45 The simplification tightrope on the pricing page24:45 When a test wins for the wrong reason27:45 Where experimentation at Box goes nextTakeaways- Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing.- Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.- Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.- Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts.- Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives.Connect with the GuestDanielle Olean LinkedIn: https://www.linkedin.com/in/dolean1/About the guest: https://www.growthbook.io/podcast/guests/danielle-oleanEpisode page: https://www.growthbook.io/podcast/episode/1-26Company Website: https://www.box.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • Dilligent explains why moving on from an experiment might cost you 07.07.2026 21min
    SummaryDan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B.Chapters00:45 Meet Dan Layfield and Diligent01:45 Two worlds of experimentation, Codecademy and Uber03:45 The trial model that lifted conversion 35%06:20 What to do with a losing experiment08:50 Two flavors of experimentation09:45 Reading forty metrics at Uber Eats13:10 Escaping the B2B feature factory16:45 Anchoring the North Star to real usage19:15 Where AI fits in research and analysisTakeaways- A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round.- Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase.- Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk.- Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them.- Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks.Connect with the GuestDan Layfield LinkedIn: https://www.linkedin.com/in/layfield/About the guest: https://www.growthbook.io/podcast/guests/daniel-layfieldEpisode page: https://www.growthbook.io/podcast/episode/1-25Company Website: https://www.diligent.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io
  • The metric Stitch Fix says every experimenter should chase 02.07.2026 20min
    SummaryIn this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs.Chapters00:00 Cold open and welcome to the show01:45 What Stitch Fix actually does04:15 Balancing AI with the human stylist05:15 From public policy to the A/B testing adrenaline rush07:15 Inside the weekly experimentation review group08:45 The AI style assistant and listening to qualitative feedback10:45 Why adoption friction beats product bugs13:45 Testing for losers and building guardrails15:45 Keep rate, successful fixes, and the holy grail metric18:15 The new platform and the promise of agentic AITakeaways- The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it.- A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.- Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping.- The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior.- Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.Connect with the GuestNick Beyler LinkedIn: https://www.linkedin.com/in/nick-beyler-381864119/About the guest: https://www.growthbook.io/podcast/guests/nick-beylerEpisode page: https://www.growthbook.io/podcast/episode/1-24Company Website: https://www.stitchfix.comSponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.Go to http://growthbook.io

Populaire dans

Ce podcast figure aussi dans les classements de podcasts de ces pays.