Writing
Sep 3, 2026CRM12 min

Why CRM Experimentation Doesn’t Scale, and How to Fix It…

Most CRM teams don’t have an experimentation problem. They have an experimentation operations problem.

Open in Medium

Most CRM teams don’t have an experimentation problem. They have an experimentation operations problem.

Almost every CRM organization who would like to streamline and optimize their communications wants to experiment more.

Test another subject line, change the audience, test different send times, run the same experiment in Germany, then France, then the UK… And then take what worked and apply it to the lifecycle journey.

On paper, none of this sounds particularly difficult. However the problem starts when its repeated across multiple markets * channels * teams.

An experiment that takes a few hours to configure isn’t much of a problem when you run five experiments per month. When you are running dozens of them across several markets, those same few hours suddenly become hundreds of hours of operational work for your CRM / customer engagement team.

And this is where experimentation starts becoming difficult to scale: not because teams are running out of ideas, but because experimentation demand starts increasing faster than the operational capacity that is required to support it.

Experimentation Costs Time… and Effort

When we talk about experimentation velocity, we usually talk about the experiments themselves.

How many tests did the team launch?

How quickly did the team reach statistical significance?

How many winning variants did the team identify?

But there is an entire operational lifecycle sitting underneath every one of those experiments.

Someone (more realistically multiple people w/ various handovers) has to:

1. define the audience

2. configure the segmentation

3. create or duplicate the campaign

4. configure the variants

5. apply naming conventions

6. configure tracking

7. validate links

8. validate personalization

9. validate audience sizes

10. configure control groups

11. perform QA

12. coordinate approvals

13. launch the experimnt

14. monitor it

15. collect the results

16. document the outcome

17. translate the learning into another campaign or lifecycle journey

A lot of this work is absolutely necessary, but very little of it is strategically interesting. Much more relevant for our discussion, it is repetitive and manual.

This distinction matters because CRM teams can easily reach a point where they spend more time actually executing experimentation than actually thinking about what they should experiment on.

That is usually the first indication that the operating model is beginning to hit a ceiling.

Now multiply that by every market

The problem becomes much more visible in decentralized CRM organizations. Imagine, for instance, a company operating across 15 or 20 markets.

A local CRM team in Market A wants to test a new onboarding incentive. They create an audience, duplicate an existing campaign, configure the experiment, add tracking, perform QA and launch it.

Market B then also wants to run something very similar.

They create the audience, duplicate an existing campaign, configure the experiment, add tracking, perform their QA and also launch it.

Then Market C does the same thing…

From the perspective of each individual market, the process is quite reasonable. But from the perspective of the organization, each market is repeatedly solving the same operational problem.

The experiments themselves might be different; for instance, the hypotheses, content, or audiences might indeed need to vary between each market’s local strategic needs, but the underlying mechanics / processes often don’t.

Those underlying mechanics are usually deterministic processes, yet in many CRM organizations they are still being executed manually, independently and repeatedly. This is where local autonomy can turn into redundant operational effort.

Manual QA becomes the next bottleneck

The more experiments a CRM team launches, the more things that team members need to check. This obvious fact comes with an important consequence.

If experimentation volume doubles while a team’s QA process remains manual, the QA workload doubles with it. Eventually the organization reaches a point where one of two things happens, either:

a) experimentation slows down because teams cannot support the operational workload, or

b) teams move faster by reducing the depth and consistency of QA.

Neither is particularly an attractive outcome for the quality and the quantity of experiments.

And the problem is not simply that manual QA takes time, but that it can also be very inconsistent (especially across markets). One person might carefully validate campaign naming, audience configuration, tracking parameters and personalization. Another might focus primarily on content and links. Someone from a different market may have created their own checklist entirely. And over time, these small differences accumulate and result in a quantifyable data problem.

Over time, two experiments that were supposed to be comparable aren’t quite measuring the same thing anymore, and the team ends up comparing apples and oranges. Ultimately, the operational problem shifts to become a data problem.

Fragmented execution eventually creates fragmented measurement

This one is a less obvious consequence of decentralized experimentation. When campaign execution varies across markets, measurement usually starts varying with it. In addition to introducing lead-time to acquire measurements, reporting itself requiring manual and market-specific work creates an additional delay between running an experiment and then translating its learnings into BAU.

The purpose of experimentation shouldn’t be to create dashboards showing which variant won. The purpose should be to generate learnings that improve the next decisions to come.

If it takes weeks for an experiment to be analyzed, documented and translated into a lifecycle campaign, it results in sunk cost for the organization. Experimentation velocity isn’t determined by how quickly you can launch the test, but by how quickly the organization can move through

Hypothesis → Experiment → Measurement → Learning → Implementation

Optimizing only the “Experiment” part of that process doesn’t really solve the underlying problem.

So what should actually be standardized?

This is where I think CRM organizations sometimes can make the wrong tradeoff. In my experience, they often choose between:

a) everything is centralized globally, which creates consistency but removes local flexibility. Or,

b) everything is decentralized, which preserves local ownership but creates duplication.

I don’t think either model is particularly good. The better question to ask is “which parts of experimentation require human judgment, and which parts are deterministic enough to standardize or automate?”

Local teams should probably continue to have autonomy over decidisions that require context, like: what should we test? which customer problem are we trying to solve? what messaging makes sense in this market? are there local commercial, cultural or legal considerations?

But there is considerably less strategic value in having every market independently decide how a campaign should be named or how experiment metadata should be generated.

A useful operating principle therefore becomes to centralize the mechanics while decentralizing the strategy.

Global teams define the framework, governance and reusable infrastructure, while local teams retain ownership of experimentation strategy, content and market-specific requirements.

We shouldn’t remove local autonomy, we should only stop requiring local teams to repeatedly perform work that doesn’t benefit from being local.

Treat experiments as configurations, not projects

Imagine a standardized experiment requesting mechanism that collects information like:

market

channel

experiment type

audience

hypothesis

campaign template

variants

primary KPI

experiment duration

Once submitted by the experiment owner, an automation layer (preferrably owned by a global automation team) can validate the request against globally defined rules. For instance, does the defined audience exist? Is the selected experiment type supported? Are the required fields populated? Does the requested campaign template exist?

If everything passes validation, the system can generate the naming conventions, generate any remaining metadata, duplicate an approved campaign shell, activate the required audience, apply tracking conventions, create an experiment record and prepare a draft for the local CRM team.

The result is a launch-ready draft where the repetitive technical setup has already been completed and validated, ready for local teams to do their final reviews and launch when desired.

Templates are more powerful than they look

Reusable blocks and/or campaign templates are infrastructural enablers when it comes to CRM automation. If your organization’s most common experiment types can be represented by approved campaign shells, you create a stable technical foundation for automating campaign creation.

The automation layer no longer needs to construct campaigns from scratch. Instead, it just duplicates something that is already known to work. This alone reduces complexity significantly.

The template becomes the governed technical implementation, and the experiment request becomes the configuration. The automation layer simply connects the two together.

This is very similar to how we approach reusable components in software engineering. You don’t want every developer rebuilding the same component every time it appears in an application. It’s merely applying software engineering principles to CRM.

Standardization should come before automation

There is an important caveat here. You cannot effectively automate a process that nobody agrees on. If five markets have five completely different definitions of how an A/B experiment should be configured, building automation around those processes will simply encode the inconsistency.

Before automating, a minimum global standard needs to be reached across the markets. That doesn’t necessarily imply that every market-specific difference will be eliminated, these edge cases will have to remain manual aspects of the newly automated process. It just means that the common denominator between everybody is identified. And for this, what better way than running a thorough discovery with all impacted stakeholders, both at the operational level as well as the strategic / senior levels.

Start smaller than you think

I am only including this section because it is a tendency I sometimes tend to suffer: scope-creep and premature optimization. If you are anything like me, everything that we have talked about so far can trigger a temptation to immediately design the new “Global Experimentation Platform” for your organization. As difficult as it might be, I would resist it.

It is better to have a working MVP tested and verified (which is also important to have strong backing from the senior leadership) first before going in this direction. Start with discovery sessions with all of your stakeholders, understand their needs, their painpoints, what can be automated and what needs to remain manual.

Then from your findings, focus as an MVP on one market (ideally the most impactful), and perhaps one channel. Automate the boring and redundant parts while keeping their wishes and needs intact. Create a simple experiment intake with field validation and automated metadata generation. Create a few approved campaign templates with the help of the local market, and introduce a template duplication process. Take care of the naming and tracking conventions. Maybe also build a small notification service that notifies the localCRM team when the draft is ready, or an error handler if something goes wrong.

Then measure what happened. Did campaign setup time decrease? Did QA issues decrease? Did the team actually adopt the workflow? Has launching experiments become faster? Most importantly, and this is something that requires more time to measure, did removing operational work result in more experimentation?

If the answer is yes, then and only then expand the scope. A few pilot markets. More experiment types. Audience activation. Automated QA validation. Experiment logging. Reporting integrations. You name it, the sky is the limit.

Eventually, you can start thinking about a genuinely global experimentation operating model. But as per product management philosophy, the platform should emerge from validated operational improvements, not the other way around.

I have seen the same pattern outside experimentation

What makes me particularly convinced by this approach is that the underlying problem isn’t unique to experimentation.

I’ve worked on a similar transformation around global CRM campaign production.

The organization was operating across 20+ markets and sending hundreds of millions of emails per month. Campaign setup and QA contained significant amounts of manual work, while technical complexity created a heavy dependency on a small number of CRM automation specialists.

Similarly to everything outlined above, the solution wasn’t to remove local CRM teams from campaign execution, but to centralize the complexity they shouldn’t have needed to manage in the first place. In fact, this whole article was formed via findings I gathered through mainly this (along with a few other similar) project[s].

I identified the common technical logic, built reusable components and dynamic content, standardized governance and naming conventions, validated the framework with pilot markets and local ambassadors, ran regression and stress tests, and only then started scaling it globally.

The technical solution was only part of the work. Training, local ambassadors, hands-on workshops, adoption tracking, documentation and eventually sunsetting legacy assets were just as important.

That experience reinforced something that I think applies to experimentation as well. Scaling CRM isn’t primarily about building more automation. It is about designing an operating model where automation, standardization and human judgment are applied to the right parts of the process.

The real metric is experimentation capacity

Ultimately, the goal isn’t automation for automation’s sake. While saving three hours of campaign setup time sounds nice, the interesting question for leadership (or the organization) is what happens to those three hours.

If they simply disappear into another operational task, you’ve made the process more efficient. If they allow the CRM team to launch another experiment, analyze another customer problem or spend more time designing better hypotheses, you’ve increased the organization’s experimentation capacity. The latter is a much more meaningful outcome.

Because at a sufficient scale, experimentation isn’t necessarily constrained by ideas. There will almost always be something else to test or try out. The constraint is the machinery required to turn those ideas into reliable experiments.

And once experimentation demand starts growing faster than that the processes can allow, hiring more people to operate the processes isn’t the answer.

Sometimes the better answer is to simply redesign the processes.

If you have made it this far, and are interested in going a level deeper, I’ve made the original case study that inspired much of this thinking available on my website. It shows what this approach looks like beyond the theory, covering how I tackled the problem for a real organization from both a technical and organizational perspective. This includes the proposed system architecture and automation workflows, phased implementation and rollout strategy, governance model, as well as the project and change management required to drive adoption and make the solution work in practice. If you want to see how the ideas explored in this article translate into an actual technical solution and implementation plan, you can find the full case study at marcmerih.ca :)