Research brief · two pages · September 2026

Can a model of one beat a model of everyone?

A pre-registered cohort study, designed and powered, waiting on an ethics process. Karthikeyan NG, independent researcher — hello@getdailyvox.com

The question

Does an emotion model adapted on a person's own labelled diary entries predict that person's later self-reported feelings better than a generic model trained on everyone — and how many entries does the adaptation take?

Consumer AI personalises by centralisation: a large model trained on everyone, tuned to an individual by sending that individual's data to it. The alternative is a model of one, built and kept on the person's own device — defensible today on ownership and privacy grounds, but with its predictive claim unproven. That claim is an empirical question with a number attached, and it can come back negative.

The longer arc is a complete personal model running on the owner's own hardware — a mirror of what someone has written, not an oracle claiming to be them. The substrate has arrived, so the open question is shifting from can it run? to how would you know it is any good? A model built to be private cannot be watched in production: no dashboard, no A/B test, and the only witness is the person being modelled, who cannot audit a model of themselves.

What already exists — no infrastructure to build

−25.2%Deployed model's trait lift on a real diary — below a do-nothing baseline
43.5%Keyword emotion layer vs an 85.0% majority baseline
r = +0.13Openness head on 1,307 writers — real, and worthless
r = +0.663The layer that does carry signal

Three of those four are the product losing. An instrument that cannot catch its own maker's model is not an instrument.

The study

Fit a per-person emotion head at K ∈ {0, 5, 10, 25} adaptation entries; grade each participant on their own later entries, primary claim at K = 25. Each contributes ~60 self-labelled entries; everything past the adaptation window is held out, capped at 35. Primary test is an exact sign-flip permutation on per-participant Δaccuracy, win-rate binomial second in a fixed sequence, ties counting as losses. The analysis runs once, after the data freezes — enforced by the harness refusing a second run, not by a promise.

ParticipantsPrimaryCompanionReading
N = 80.590.83Underpowered; "inconclusive" is a likely honest outcome
N = 100.92Companion test becomes decisive
N = 150.75First size tolerating three participants losing

Anti-fabrication gate. A leave-one-participant-out donor arm is pre-registered as an interpretive gate: adapt on other participants, test on this one. If donor adaptation matches personalised adaptation, the personalisation claim is withdrawn and the result is published as spoken-register domain adaptation instead. The study is publishable on either branch.

Why it cannot run today

Three blockers, none technical — which is the point. They are what one unaffiliated person cannot supply.

One item is open rather than blocked: the pre-registration is still a draft. It becomes unamendable once frozen, so it stays unfrozen until someone qualified has attacked it — which is the highest-leverage first contribution, and has to happen first.

Division of labour

The collaborating group supplies

  • An ethics process and institutional standing.
  • Adversarial review of the protocol before it freezes.
  • Recruitment and consent of 8–15 participants.
  • Literature depth in affective computing and personalisation.

I supply

  • The app, the labelling instrument, the export tooling.
  • The harness, the power analysis, the analysis code.
  • All engineering and data plumbing.
  • Drafting, revision, submission mechanics.

Co-authorship agreed in writing before anyone starts. The study is the right size for a master's or final-year project, with a publishable result on every branch — including the null.

Where it publishes — probable targets, not a plan

A Registered Report fits best, and is available now. "Inconclusive" is the modal outcome here and every conference rewards results; a Registered Report is accepted on the design before data exists, making that outcome publishable by construction. PCI Registered Reports is free and rolling, states it "sets no minimum requirement for statistical power," and its Stage-1 acceptance is honoured by 41 journals without further review — including Royal Society Open Science and PeerJ Computer Science. No speech, NLP or ACM HCI venue has such a track.

VenueDeadlineNote
UMAP 2027 · Chicago22 / 29 Jan 2027"Does adapting to this person beat the population model" is UMAP's own question
IMWUT (UbiComp)1 Feb 2027On-device sensing home turf; journal R&R rather than binary reject
MobileHCI 202710 / 17 Feb 2027Friendliest pool — in-the-wild deployment at this scale is a normal contribution
PoPETs 2027 · Issue 428 Feb 2027Strongest artifact evaluation; needs a privacy-measurement framing
IEEE ASLI 202721 Apr 2027Merged ASRU + SLT successor; the latest date, buying ten more weeks of collection

On whether this cohort size publishes at all: JMIR Formative Research has run a person-specific study at N = 5 and an n-of-1 per-person modelling study. Rolling alternatives: IEEE TAC, ACM TiiS, ACM HEALTH.

Risks, stated

Ethics, settled before anyone asks: informed consent, no compensation, data kept local to the participant's machine, deletion honoured immediately, withdrawal any time before publication, aggregate numerics only in anything published. Where a collaborating institution's process is stricter, its process governs.

getdailyvox.com/research-statement (full version) · /research (programme) · /paper/measuring-a-model-of-one.pdf — the pre-registration draft is available on request, and the most useful first response to it is suspicion.