The takeaway

Density is context, not a quality verdict. State the employee population, time window, record definitions, and distribution behind each number.

A company-wide total hides the shape of the data

Two businesses can export the same number of messages and represent very different kinds of work. One archive may cover a small active team over several months. Another may include years of automated alerts with little human decision-making.

Density adds a useful question: how much relevant context exists for the people and period represented? For a buyer trying to reconstruct work, that context can be more informative than headcount alone.

Review density alongside cross-system linkage and workflow completeness. A dense archive with no reliable relationships may still be difficult to use.

Choose a denominator you can explain

A useful starting unit is an active employee-month, defined for the agreed sample. Explain whether activity means employment during the period, participation in the included systems, or an observable event. Different definitions answer different questions.

Illustrative example: 12,000 eligible email events across 100 defined active employee-months gives a mean of 120 events per employee-month. That does not tell you the median, the number received versus sent, or how activity is distributed across roles.

Do not multiply current headcount by a historical period and assume that describes the population. Account for joiners, departures, contractors, partial months, and systems used by only part of the business.

Count each type of context deliberately

Questions behind common density measures
Record typeDefine before counting
EmailSent events, received events, unique messages, or threads?
ChatHuman messages, automated notifications, or both?
FilesUnique files, revisions, attachments, or access events?
CommentsUnique comments, repeated exports, or edit history?
CalendarUnique events or one attendance record per person?

A message copied into several mailboxes should not silently become several independent examples. Likewise, a legitimate file revision is not necessarily an accidental duplicate. Preserve the unit that matters to the task and document the counting rule.

Show a distribution, not just an average

Report each record type separately, with its median and a useful spread such as lower and upper quartiles. Segment by relevant role, system, or period when the scope supports it. A handful of highly active accounts can otherwise dominate the mean.

Separate “no activity observed” from “source not exported.” A quiet month and a missing month should not look identical. Show which employees have consistent identities across the included systems; our identity guide explains the checks.

Do not add separately computed medians and label the sum a median total. If a combined measure is useful, calculate it at the individual employee-month level first.

Ask whether the volume helps the intended task

Google Research’s work on deduplicating language-model training data reports benefits from removing duplication in the studied training setup. It does not mean every repeated operational event should be deleted: repetitions can also document an escalation or revision.

Our recommendation is to separate duplicated exports from meaningful repeated work. Then inspect whether the sample contains decisions, exceptions, and outcomes, rather than using a high message count as a proxy for quality.

Include the resulting definitions and distributions in a buyer-readable manifest. This makes the count reproducible and its limits visible.

Use density to improve the conversation

For a lab, density can help identify which samples deserve closer task-level review. For a business, it can explain operational depth without making an unsupported valuation claim.

Discuss quality ranking with Clear Harness, our partner, and bring the record definitions with the totals. To explore your company’s separate licensing opportunity, take the Troveo assessment (referral link; we may be compensated).

Sources & further reading

Sources checked October 9, 2026.

About Ohio Training Data

We help businesses explore data licensing through Troveo and connect them with Clear Harness, our partner for business-data quality ranking. Have a question about the next step? Get in touch.