Updated August 2026. Originally published May 2024.

Between January 2017 and January 2018 my team ran 59 A/B tests.

We won 26 of them. We lost 31. Two were inconclusive.

That is a 44 percent win rate. I lost more tests than I won across a full year, and I have the tracker to prove it.

I am opening with that because it is the number nobody in this field publishes. Every article about conversion rate optimization is written as though the author wins, and the reader is left assuming that a good practitioner picks winners. A good practitioner does not pick winners. A good practitioner runs enough tests that the winners pay for the losers.

Why I am showing you a record from 2017

That year is deliberate. It is the last period where the entire discipline was visible from the outside.

Every hypothesis was written down before it ran. Every result was logged by hand. Every conclusion was argued over in a meeting, because there was no other way to reach one. Nothing was hidden inside a tool that summarised it for us.

That visibility is what lets me show you how this work has changed rather than simply assert it. The reasoning that produced those 59 tests is the same reasoning I use today. What has changed, and changed enormously, is what it costs to run that reasoning and how long you wait to find out whether you were right.

So read the next few sections as the old cost of the method, and the last ones as the current cost of the same method.

The part that matters more than the win rate

Over that same year, with a losing record, conversion went up across every product group we were responsible for:

  • Certificate products: 7.43 percent to 9.38 percent
  • Security products: 14.25 percent to 16.94 percent
  • Enterprise products: 17.82 percent to 19.12 percent

Losing 31 tests was not failure. It was the cost of finding 26 things that worked. The losses were cheap, because a losing variation gets switched off. The wins compound, because they stay on.

That is the entire economic argument for testing, and it only becomes obvious when you keep the record.

What the setup actually was

Three people, full time, on site, doing nothing else.

The tooling was Optimizely for the tests, Google Analytics and Hotjar for behaviour, the CRM for what actually closed, and Excel for everything else. There was no dashboard that told us what was happening. Results were pulled, pasted, tracked by hand, and argued about in a spreadsheet that eventually ran to thousands of rows.

The three of them read the traffic data and built the hypotheses. That was the job, and it was skilled work. What surrounded it was not. Pulling numbers, maintaining the tracker, chasing implementation, formatting reports. The thinking was theirs. The overhead was the price of doing it in 2017.

Keep that in mind for the last section, because the thinking has not changed and the overhead has almost entirely gone.

How a single test actually ran

This is from the process document we worked to at the time, and two steps in it are still the difference between testing and guessing.

  • Hypothesis first, from analytics. Where to test, what to test, and why. The why is the part people skip.
  • Sample size and duration calculated before launch, with a confidence interval, using a calculator rather than intuition. You decide when the test ends before you can see which way it is going.
  • Wireframes, approval, then a developer to place the testing snippet in the page and implement the variation.
  • Primary and secondary goals set explicitly in the tool, with tracking parameters so results could be matched back to the CRM.
  • Results validated in a second system. The testing tool told us the conversion difference. Analytics and the CRM told us whether it produced revenue, compared week against following week.
  • Canonical tags on any split URL test, so running two live versions of a page did not damage search rankings.
  • Losers got a fresh hypothesis, not a shrug. The page went back into the queue.

Two of those deserve emphasis, because they are what most people skip.

Never trust the testing tool alone. Optimizely will happily report a winner. Whether that winner produced money is a separate question, answered in a different system. We checked, and sometimes the answer disagreed.

A split URL test without a canonical tag is an SEO problem. You are serving two versions of the same page to crawlers. I have written more about how that structure should work in internal linking for SEO.

The one principle I would keep if I could only keep one

Write in the words your buyer uses. Not the words you would use.

The term I use for this is the buying trance, which I take from Joe Vitale’s book of that name, so credit where it is due. The practical version is simple and most brands get it backwards.

A brand does not get to decide how its offering is described. The persona decides. Your job is to find their language and use it exactly, including the parts that sound wrong to you.

The test that taught me this

We were testing a landing page for an MSP audience. I built what I thought was the better version: proper design, considered layout, and copy that described speed in precise technical terms, because the product was technical and the buyers were technical.

It lost. It lost to a plain page on a white background with simple wording.

The phrase that beat me was “breakneck fast”.

I had never heard it before. It was not a term I would have written, and if a copywriter had handed it to me I would probably have edited it out. It was simply how those buyers talked to each other, and using their slang beat describing the same thing accurately in mine.

I am working from memory on that one rather than from the tracker, so treat the phrase as remembered and the lesson as the point.

The losses taught me more than the wins

Three from that year’s record, all of them things I was confident about:

  • A signup page redesign lost 21 percent. It was logged in the tracker as a radical design test, which tells you how ambitious it was. The plain control beat it.
  • Adding proximity between the copy and the call to action lost 42 percent. That is a textbook best practice. It is in every CRO course, including the one I was certified on. On that page it was wrong.
  • A home page first fold rebuild lost 49 percent. New layout, new hierarchy, better by every design argument anyone made in the room.

Notice what those have in common. They are all cases where the thing that lost was the thing an expert would recommend. Best practice is a hypothesis, not an answer, and the only way to find out is to run it.

The result that ended best practice for me

In 2020 I ran the same test across nine different MSP businesses. Same hypothesis, same change: a new banner headline and call to action. Similar companies, same industry, same buyer.

The results were not similar. They were opposite.

  • Four sites improved substantially, the strongest by roughly 150 percent
  • Others fell by 42, 55 and 72 percent
  • One dropped to zero conversions on the variation

Same test. Same vertical. Opposite outcomes.

If a change that works on four businesses destroys conversion on four comparable ones, then there is no such thing as a portable answer in this discipline. There is only your audience, your page, and what happens when you actually check.

This is also why I am sceptical of case studies that report a single number with no context. The number is real. It just does not transfer to you.

Before you test anything, understand who is already there

Most people start by choosing a change. Start instead with the people already on the page.

  • What did they search for or click to arrive here
  • What do they already believe by the time they land
  • What are they afraid of, and what would make them leave
  • What are they comparing you against, in their words

A hypothesis built from that has a chance. A hypothesis built from a list of tactics is a coin flip with extra steps. If you have not defined what you are actually promising, start with the value proposition before you test anything on the page.

How I run CRO in 2026

The method has not changed. The overhead has collapsed, and so has the number of things I have to assume.

Everything lives in one folder

What used to sit across spreadsheets, a certification binder and my own memory now sits in a single folder the system can read:

  • My actual test results going back years, wins and losses both
  • My conversion optimization certification material
  • The AIDA framework, and my own documentation on priming
  • The business’s internal and external material, so the context is real rather than generic

That last point is the one people miss. A general model can tell you what usually works. A system holding your own losing tests can tell you what already did not work on your pages, for your buyers.

The frameworks stopped being theory

AIDA, priming, proximity, message hierarchy. In 2017 these were things you applied by hand and then tested one at a time, and each test cost weeks of calendar time between design, approval, implementation and sample collection.

Serial testing of that kind puts a hard ceiling on how much you can learn in a year. Fifty nine tests was a lot. It was also, in practice, fifty nine questions answered in twelve months by three full time people.

Now those frameworks sit in the same folder as the results, so they are applied to the analysis rather than recalled imperfectly during it. They are not a checklist I try to remember. They are part of how the material gets read.

The real gain is fewer assumptions

This is the part I would emphasise if you take one thing from this section.

In 2017 a hypothesis carried a stack of assumptions underneath it, and checking any one of them was expensive. So we did not check them. We bundled them into the test, ran it for three weeks, and if it lost we usually could not tell which assumption had been wrong. That is a large part of why the loss rate was 53 percent.

Most of those assumptions can now be checked before anything is built. What the traffic already does, how the segments differ, whether the page even gets read to the point you are changing, whether a similar change has already failed somewhere in my own history. The test still decides, but it starts from a much better position.

Fewer assumptions going in means fewer tests wasted on questions I could have answered first.

What the loop looks like now

If you want to copy the process rather than the tooling:

  • Put your own history in one place first. Past tests, past results, your positioning material, your buyer research. Do this before you ask anything of an AI, because without it you get generic advice dressed up as analysis.
  • Ask it to interrogate the data, not to suggest tactics. Tactics are free and worthless. Ask what the traffic is doing, where the drop is, which segment behaves differently, and make it show you where in the data it got that.
  • Check the hypothesis against your own losses before building it. This is the step nobody has, because nobody keeps the record.
  • Implement it yourself. I make changes over SSH rather than waiting on a queue. The variation exists in minutes, not next sprint.
  • Watch the metrics live rather than exporting them and pasting them into a spreadsheet on Friday.
  • Keep logging everything. The record is still the asset. It is now also the input.

Faster testing, on better hypotheses, with results you see immediately, feeding a history that makes the next hypothesis better. That compounding is the actual change, more than any individual capability.

What has not changed

The test still decides. Not me, and not the model.

AI removes assumptions and it removes waiting. It does not remove being wrong.

Everything in the losses section above was produced by an experienced team with good reasoning, and best practice still lost by 42 percent. A faster wrong answer is still a wrong answer.

The discipline that protects you is the same as it was in 2017. State the hypothesis. Decide the sample size and the duration before you start. Let the result stand even when you dislike it.

Look at that 2017 checklist again. Wireframes, approval, a developer to place the snippet, a developer to implement the variation, then manual validation across three systems. Every one of those was a queue, and a test could sit in one for a week.

The loop is the same loop. What has gone is the waiting between its steps, and much of the guessing inside them.

This is the same pattern I found when I documented an SDR’s day and priced the gaps, which I wrote up in how I moved a marketing team onto AI. The work that disappears is never the thinking. It is the connective tissue around it.

The prediction I made, and where it stands

In 2024 I predicted that bots would move past answering questions to analysing behaviour, forming their own hypotheses from stored data, and eventually executing those hypotheses through content placement and real time content replacement based on what a visitor did on the previous page.

The capability arrived. The adoption did not.

I have not seen anyone doing this properly on a live site. I got there myself, in the sense that the system I work with now does form hypotheses from my own history and act on them. But this is not how the industry operates today, and I was wrong about how fast that would spread.

Predicting the technology turned out to be easier than predicting whether people would use it. That gap is worth more attention than it gets, and it is closely related to why I now argue that pages need to be legible to systems as well as people, which I set out in the piece on getting cited by AI.

If you are starting on this

  • Keep the record. Every test, won lost or inconclusive. Without it you will remember only your wins and learn nothing.
  • Expect to lose more than half. If you are winning most of your tests you are testing things that are too safe to matter.
  • Steal your buyer’s words. Read how they actually talk and use their phrasing, even when it sounds wrong to you.
  • Treat best practice as a hypothesis. Proximity, hierarchy, social proof. All of these are worth testing and none of them are worth assuming.
  • Do not port results between businesses. Not even between businesses that look identical.
  • Measure the cumulative effect, not the individual test. A losing year of tests can still be a winning year of conversion.

Everything above came out of a spreadsheet I still have. That is the only reason I can tell you the win rate instead of just telling you it went well. For how I think about proving marketing performance more broadly, see measuring marketing ROI in high-touch sales.

This article was substantially rewritten in August 2026. The original May 2024 version set out the principles and process without the underlying results. It has been rebuilt around the actual test record from 2017 and 2018, the process document my team worked to at the time, and the multi-site MSP tests from 2020. Company, client and product identities are deliberately withheld, and the figures are reported in aggregate. The ecommerce funnel checklist from the original has been removed as out of scope. The prediction from the original is preserved and graded, and a section on how I run this work today has been added.

Categories: Blog

Ugur Gulaydin

Vice President of Marketing at Corporate Technologies, a managed IT services provider working with small businesses from 21 locations across 18 states. Over a decade in B2B demand generation across cybersecurity, managed IT services, home automation and cloud security, including more than 2,000 conversion tests and over a thousand inbound campaigns. Everything on this blog is written from work I have actually done, not from what the playbooks say should work. More about me · LinkedIn