Skip to content

Cohort retention table

See whether users come back over time. Paste a raw event list or a ready cohort matrix; the tool assigns cohorts and periods and shows the retention percentage triangle, the weighted average retention curve and the differences between cohorts.

Free tool · Data analytics
Input type

Each row is a cohort: the cohort label, then the active counts for period 0, 1, 2 … Period 0 is the cohort size. Leave periods not yet observed empty (rows may differ in length). Counts must be whole numbers without thousands separators. The first row may be a header.

The data is read and calculated in your browser; it is not sent anywhere.

Let us find together, in your data, the factors that would raise your customer or user retention.

Request a call

01

How to use

  1. A

    Choose the input type: a ready cohort matrix or a raw event list (user; first date; activity date). You can also try the sample data.

  2. B

    For an event list choose the period length (day, week, month); the tool assigns cohorts and periods itself.

  3. C

    Read the retention percentage triangle, the weighted average curve, the period-1 loss and the cohort comparison; download the table as CSV.

02

How is a cohort retention table read?

A cohort is a group of users who started in the same period (for example those who signed up in January). In the table rows are cohorts and columns are the periods since the start; a cell shows the share of that cohort still active in that period. Period 0 is the period the cohort started in; it is 100% if the first activity date defined the cohort.

The table is a triangle: periods that newer cohorts have not lived through yet are empty. Going down a column shows whether newer cohorts retain better or worse than older ones (the effect of a product or campaign change).

03

Weighted average curve and period-1 loss

For each period the weighted average divides the total active count of the cohorts that have reached it by their total size, so larger cohorts carry more weight. The first drop of the curve (the period-1 loss) is the biggest loss in most products and shows the quality of the first experience.

If the curve flattens over time there is a lasting core of users; if it keeps falling the product may not be sticking. The last points of the curve are computed from few cohorts and are more volatile; the tool shows how many cohorts contribute at each point.

04

How is the cohort assigned?

In an event list a user's cohort is the day, week (starting Monday) or month of their first date; if a user has several first dates the earliest is taken. The period of an activity is the difference in periods between the activity date and the cohort period. Several activities in one period count once: the active count is the number of distinct users with at least one activity in that period.

Dates are read as written, in UTC; there is no time-zone conversion, so events close to midnight may fall on a different day in the user's local time. Excel's internal serial numbers (such as 45123) are not recognised as dates; export the cells formatted as dates.

FAQ

How should I choose between day, week and month?
By the natural usage rhythm of the product. Day or week suits daily-use apps, month suits monthly billing or subscriptions. Very short periods produce many cohorts and noisy percentages; at most 120 cohorts are supported.
Why is period 0 not always 100%?
The cohort is set by the first date (for example sign-up) and the activity may not fall in the same period. If fewer users are active in period 0 than the cohort size, period 0 is below 100%. In matrix input the period-0 count is the cohort size and is 100%.
Why is the last column always low?
The last observed period is usually not over yet; only the activity so far is counted in it. That is missing data, not a loss. Look at completed periods for decisions.
What happens if an activity date is before the first date?
A negative period is meaningless; those events are not counted and the warning shows how many were skipped. It often means the first-date column is defined wrongly or dates from different systems are mixed.
Why does the weighted average differ from the average of the cohorts?
The weighted average goes by the number of users: a large cohort counts for more. A plain average of the cohort percentages counts every cohort equally; an extreme value in a small cohort moves it too much. The weighted average is usually more stable.
Does one cohort doing better show the cause?
No. The difference may come from a campaign, season, price or product change, and in small cohorts it is chance. A cohort comparison raises a question; it does not prove the answer. Test the cause with an experiment or more data.

Let us find what raises retention

Let us find out which behaviours in your user or customer data predict long-term retention and turn that into a measurable improvement plan.