Comparison table with rows, controls, and labeled cells. Platform comparisons data: what the experienced already know, updated for 2027
Image: Content Publishing Systems

Strategy

Platform comparisons data: what the experienced already know, updated for 2027

Platform comparisons data comes from one experiment, not an assembled table: the four controls, the three kinds of cell, and the record to keep.

Most comparison tables are assembled, not measured. One cell comes from a pricing page read in March, another from a September test run, a third from the writer's memory. Each cell may be accurate.

The table is still not a comparison. A comparison is a claim about a difference. Two measurements taken months apart under different conditions do not show a difference between the things.

The fix is to treat a comparison as one experiment with a start and an end, rather than as a document that gets filled in. This page is about how to run it that way, what has to be held constant, and what to record so that somebody else could repeat it.

What to take away

  • Every cell in a row must be produced the same way, in the same window, by the same instrument. A row assembled from different moments measures the moments.
  • Hold the account, the content, the window, and the reader constant. Those four are the controls, and without them a result is uninterpretable.
  • Publish the protocol beside the table. A comparison without its method is a set of assertions in a grid.

The unit of work is a row, not a cell

Think of each criterion as a small experiment across all candidates, run once, start to finish. That framing has an immediate consequence: you do not fill in a table left to right by service, you fill it in row by row, doing the same thing on every candidate before moving on.

Row-by-row vs service-by-service

Row by row

Order
All candidates per criterion
Risk
No time gradient
Example
Same window for all
Fix
State both dates in row

Service by service

Order
All criteria per service
Risk
Unnoticed time gradient
Example
First before pricing change
Fix
Footnote nobody reads

Doing it the other way is how a table acquires an unnoticed time gradient. The first service was measured before a pricing change and the last one after it, and the resulting difference is real, dated, and about the calendar rather than about the services.

Where a row genuinely cannot be completed in one window, say so in the row rather than in a footnote nobody reads, and give both dates.

The four controls

An experiment needs things held still so that the one thing you are varying is the one you are studying. The design of an experiment is mostly a question of what you refuse to let vary, and in a platform comparison there are four candidates.

Four controls to hold still

  • Accountsame age, activity, region, follows
  • Contentproduced before the run, not during
  • Windowsame length, same time, simultaneous
  • Readerone job, requirements fixed before scoring
  • Name any confounder in the protocol

The account. Same age, same prior activity, same region, same followed accounts, created the same way on the same day. An established account and a new one are different instruments and will disagree about the same platform.

The content. The same material, adapted only as much as each platform requires, produced before the run rather than during it. Improving as you go makes the last candidate look best.

The window. The same length, at the same time, ideally simultaneous. Sequential runs pick up seasonality and whatever else happened in between.

The reader. One stated job, one stated set of requirements, fixed before scoring. Changing who the comparison is for halfway through is the commonest way a conclusion gets reverse-engineered.

Anything you cannot hold still is a confounder, and the honest move is to name it in the protocol rather than to hope it averaged out. It did not average out; there is one run.

What each kind of cell is, and how to mark it

Three kinds, and mixing them without labels is what makes most published comparisons unusable.

Three kinds of cell

Measured

Example
Export contents
Record
Result, method, date
Evidence
Your own test
Mark
Kind column

Quoted

Example
Pricing page listing
Record
Verbatim quote, date
Evidence
Operator's description
Mark
Kind column

Not established

Example
Ranking weights
Record
Sought, not found
Evidence
Document classes searched
Mark
Kind column
KindExampleHow to record it
MeasuredWhat the export contained; whether a non-connection could message youResult, method version, date, artifact
QuotedWhat the pricing page listed; what the help center said a feature doesVerbatim quote, document, its date, your reading date
Not establishedHow ranking weights a post; how many accounts are realSought and not found, with the document classes searched

Give the table a column for the kind, or use a mark, but do not leave it implicit. A reader deciding on the strength of a quoted cell is relying on the operator's description of its own behavior, a different level of evidence from a test, and they are entitled to know which they have.

This is the same three-tier separation the notes on collecting data about a service apply to a single-service dataset. It matters more here because a comparison invites a conclusion.

The control the category makes hard

There is one thing you cannot hold still and cannot measure: the audience. Two platforms with the same features and different people produce different outcomes, and the difference is not a property of either platform.

Does the criterion survive?

Is the criterion structural or audience-dependent?

Yes

keep: export, portability, messaging, visibility, moderation

No

drop: reach and other audience outcomes

There is no control available for it. You cannot run the same audience through both. What you can do is say, in the protocol, that outcomes are contaminated by audience composition. The comparison is therefore only meaningful on structural criteria and on things you did to the platform rather than things the platform did for you.

That restriction eliminates a lot of the criteria people want and it is the honest boundary of the exercise. Everything about export, portability, messaging rules, visibility, and moderation survives it. Everything about reach does not.

The record to keep

One row per observation, exactly as everywhere else on this site, with the criterion, the candidate, the value, the kind, the method version, the window, and the artifact. The current table is a query over that record rather than a document you edit, which is what lets a partial re-test slot in without disturbing anything else.

Fields for every observation row

  • Criterion
  • Candidate
  • Value
  • Kind
  • Method version
  • Window
  • Artifact

Keep the artifacts. A saved export file, a screenshot of a pricing page, the raw file from a dashboard pull: these settle a later disagreement and cost almost nothing to store.

The routine for holding your own copy of the numbers is the same habit applied to a single service. A comparison needs it more, because two parties will eventually dispute a cell.

Record refusals too. A candidate that could not be tested on a criterion, because the capability sits behind a tier you do not have, is a cell with a reason rather than a blank.

Publishing it

Put the protocol on the page, above or beside the table: the accounts, the window, the content, the job, the method version, and the confounders you could not control. It is half a screen and it is what turns the table from an assertion into a result.

What to publish with the table

  • Accounts used
  • Window
  • Content
  • Job
  • Method version
  • Uncontrolled confounders

State the conclusion in the weakest form the evidence supports. "On these criteria, in this window, with these accounts, for this job, the two differ in the following two ways" is a sentence you can defend in a year.

Anything stronger claims about platforms in general. One run cannot support that. The six-axis method deliberately avoids this by comparing structures, not outcomes.

Where the comparison will be repeated, freeze all of this and version it, because a series is only a series if the instrument stayed the same. The rules for keeping a repeated comparison comparable set out what happens when it does not.

Common questions

Is a single run worth publishing?

Yes, labeled as one, with its dates. A single well-documented run beats an undated assembly of cells from a year of reading, and it can become the first point of a series later.

What if I only have time to test one criterion properly?

Test that one and quote the rest, marked as quoted. A table with one measured row and six quoted ones, honestly labeled, is more useful than seven rows of unmarked mixture.

How do I compare something I cannot test, like moderation?

You do not test it; you quote the published rules and record any observed enforcement with dates as separate evidence. The gap between the two is the finding, and it belongs in the record rather than in a score.

Does this apply to a personal decision between two services?

In miniature, and it is worth it. Same content on both, same fortnight, same measurement, decided in advance. That is an afternoon of setup and it answers the question for you rather than for whoever wrote the comparison you were about to trust.

More in Strategy

Latest from Practice Desk