← Back

Methodology

How these assessments work

Written to be checkable. Where a number comes from, which items are borrowed and which are ours, and the specific things this cannot tell you.

What is borrowed and what is written here

This distinction matters more than any other on the page, so it comes first. Four instruments are reproduced word for word, because their scoring thresholds are only meaningful against the wording they were validated on:

  • PHQ-9— depression screening (Kroenke, Spitzer & Williams, 2001)
  • GAD-7 — anxiety screening (Spitzer et al., 2006)
  • PSS-10— perceived stress (Cohen, Kamarck & Mermelstein, 1983)
  • SWLS — life satisfaction (Diener et al., 1985)

Everything else is written here, built on a published framework but not reproducing a published instrument. The Big Five items follow the OCEAN model; the Dark Triad items follow the three-factor structure of the Short Dark Triad (Jones & Paulhus, 2014); the self-esteem scale extends the Rosenberg construct with items on self-criticism and self-compassion that Rosenberg does not separate. Those are our items on their scaffolding — which means the framework is validated and our specific wording is not. Each assessment states which of the two it is on its own page.

Why there are no percentiles

No score on this site is a percentile, and none says you are ahead of or behind other people. A percentile requires a reference sample that is representative of the population you want to compare against, collected under comparable conditions. We do not have one, and inventing a number that looks like one would be the fastest way to make everything else here worthless.

What you get instead is your position on the scale's own range: 0% is the lowest answer available on every item, 100% the highest. Where a published instrument has validated severity thresholds — the four above — those are used and named as what they are. The behavioral tasks are anchored to a property of the task itself rather than to other players, which is spelled out on each task's page.

How a score is calculated

Every item belongs to exactly one sub-scale. Reverse-keyed items are flipped within their own range before averaging, so agreeing with an item written in the opposite direction counts correctly rather than cancelling out. A sub-scale score is the mean of its items mapped onto 0–100%, and the item count and raw mean are shown next to each one so you can see how much evidence sits behind it — a facet resting on three items is labelled as such rather than presented with the same confidence as one resting on ten.

Sub-scales are reported separately rather than summed into one number wherever summing would destroy the information. Attachment is scored on two independent dimensions because anxious and avoidant attachment are not two ends of one line. The five love languages are scored as five channels, because averaging them produces roughly the same middling figure for everyone. Some scales are charted inversely, where a high score is a concern rather than an achievement, and those are marked.

How the questions are written

Two styles, chosen by what is being measured. Traits that show up as behavior — the Big Five, the Dark Triad, risk-taking, decision style — are asked as concrete situations, because a situation is easier to recognize than an abstraction and harder to answer aspirationally. Things that are internal states — attachment, self-esteem, resilience, mindfulness — are asked as plain first-person statements, since wrapping a feeling in a scene means inventing a scene the reader may never have been in.

Items are kept to one idea each. An item that describes a situation and then explains what the situation means is really two questions, and someone who recognizes the first half but not the second has no correct answer available.

What this cannot tell you

Self-report measures what you are willing and able to say about yourself. Both halves are real limits: some things you would rather not put in writing, and some you cannot see from the inside. Self-reported and behavioral measures of the same trait correlate far less than most people expect — for risk-taking, correlations around r ≈ .2–.3 are typical (Frey et al., 2017). That gap is why the behavioral tasks exist here alongside the questionnaires, and comparing the two is more informative than either alone.

Nothing here is a diagnosis, and the screening instruments in particular measure how recent weeks have been rather than what you are like. A questionnaire cannot tell you why you are the way you are, and a score that moves between two takes may reflect a changed week rather than a changed person. These are useful as a mirror and a starting point for a conversation, including with a professional — not as a verdict.