What survives when information becomes data
A question under review: how much of the meaning in a body of information survives translation into something computable, and can that loss be measured rather than assumed?
Research question
Turning information into data is never lossless. Can we characterise what a given representation keeps and what it discards — and put a number on the difference?
Overview
A conversation becomes a transcript. A transcript becomes tokens. Tokens become vectors. Something is discarded at every step, usually silently, and the analysis downstream has no way to know what it lost.
Lattice is, for now, a question rather than a programme of work. Before committing effort we want to know whether the loss introduced by a representation can be characterised in a way that generalises — or whether it is only ever answerable case by case.
Approach
- Sketch small, well-understood cases where the intended meaning is known by construction, so representation loss can be quantified rather than debated.
- Survey where existing fields already answer this — information theory, compression, measurement theory, linguistics — before inventing anything.
- Decide honestly whether the question is tractable at a useful scale, and close the initiative if it is not.
Log
June 10, 2026 · Concept logged
Opened as an exploratory concept. Reading and scoping; no commitment to a full initiative.
Findings to date
Exploratory. Nothing to report — this is a question under review, not a programme of work.