Lens · Research
Every instrument in the battery, the way your results get synthesized into a profile, and the guidance you get on college and major fit are all grounded in published, peer-reviewed research — not opinion. This page lays out what that research actually says, and how we've used it.
The 9 instruments in the Lens battery — plus a separate step to add your real academic data (GPA and SAT/ACT), which is data rather than a psychometric assessment — are drawn from peer-reviewed psychological research. Each has been independently validated and is freely available for non-commercial use (or, for the Leadership Styles Inventory, adapted from an academic scale rather than a paid proprietary one). Together they cover personality, character, values, motivation, risk tolerance, passion, and leadership style — producing a richer picture of a person than any single assessment could. Two additional cognitive tasks (Matrix Reasoning and 3D Rotation) and the Remote Associates Test were retired from the required battery — see the "Currently On Hold / Retired" section below for why, and for why we still document them here.
The Big Five emerged from the lexical hypothesis — the idea that the most important personality differences in human life become encoded in language over time. Researchers catalogued personality-describing words and found they consistently clustered into five broad factors. Paul Costa and Robert McCrae formalised the model in their NEO-PI inventory (1985); Oliver John, Naomi Donahue, and Richard Kentle condensed it into the 44-item Big Five Inventory (BFI) in 1991. The model has since been replicated across 50+ countries and is the most widely used personality framework in academic psychology.
Each factor is a spectrum. Neuroticism is sometimes labelled Emotional Stability (reversed).
44 statements rated on a 1–5 Likert scale (strongly disagree → strongly agree). Roughly half the items are reverse-scored. Each of the five trait scores is the mean of its 8–10 items, giving a score from 1.0 to 5.0.
Angela Duckworth, Christopher Peterson, Michael Matthews, and Dennis Kelly introduced grit in a landmark 2007 paper in the Journal of Personality and Social Psychology. They defined it as "perseverance and passion for long-term goals" — and found it predicted outcomes (West Point retention, National Spelling Bee performance, GPA) beyond what IQ or conscientiousness alone could explain. The original 17-item scale was later condensed to the 8-item Grit-S, which is used here.
Perseverance reflects working hard through setbacks. Consistency reflects maintaining focus on goals over years rather than shifting passions.
8 items on a 1–5 scale. Some items reverse-scored. Overall grit score = mean of all 8 items (1.0–5.0). Subscale scores = means of the 4 perseverance items and 4 consistency items separately.
Ralf Schwarzer and Matthias Jerusalem developed the GSE in German in 1979, building directly on Albert Bandura's foundational theory of self-efficacy — the belief that one can execute the behaviours needed to produce specific outcomes. They adapted and validated the English version in 1995; the scale has since been translated into 33 languages and used in thousands of studies. Unlike domain-specific efficacy measures, the GSE captures a generalised sense of competence applicable across novel and stressful situations.
A person's belief in their ability to handle difficult tasks, cope with setbacks, and persist through obstacles — regardless of the specific domain. Predicts academic performance, health behaviour, and job satisfaction.
10 items on a 1–4 scale (not at all true → exactly true). No reverse-scored items. Score = mean of all items (1.0–4.0), or equivalently the sum (10–40).
John Holland developed his theory of vocational personalities and work environments over four decades, with major works in 1959, 1973, and his final revision in 1997. His central claim: people can be described by their resemblance to six personality types, and work environments can be similarly classified. Person-environment fit — the match between a person's type and their work environment — predicts satisfaction, stability, and achievement. The six types form a hexagonal model (the RIASEC hexagon), where adjacent types share more in common than opposite types.
A person's top 2–3 types form their Holland Code (e.g. "IAS"), which maps to compatible careers and study paths.
Ratings of interest in activities, occupations, and self-assessed competencies on a 1–5 scale. Score per type = mean of its items. Profile of six scores; the highest indicate dominant interests.
Shalom Schwartz proposed his Theory of Basic Human Values in 1992, identifying 10 motivationally distinct values that appear to be universal across cultures — validated in 80+ countries. The PVQ-40 (2001) operationalises this theory with an indirect measurement approach: instead of rating values directly (which can inflate scores due to social desirability), respondents read 40 short portraits of a person and rate "how much like me is this person?" This reduces acquiescence bias.
40 portraits rated on a 1–6 scale (not like me at all → very much like me). Each value score = mean of its 3–4 portrait items. Scores are ipsatized (centred on each person's own mean) to control for individual differences in scale use.
Self-Determination Theory (SDT), developed by Edward Deci and Richard Ryan from the 1970s onward, proposes that three universal psychological needs must be satisfied for well-being, growth, and intrinsic motivation. Bart Chen and colleagues published the BPNSFS in 2015 as a theoretically grounded instrument that captures both need satisfaction (which promotes flourishing) and need frustration (which actively produces ill-being) — an important distinction earlier scales missed.
24 items on a 1–5 scale (completely untrue → completely true). Six subscales of 4 items each. High satisfaction scores predict well-being, vitality, and intrinsic motivation. High frustration scores predict anxiety, depression, and amotivation independently.
Most risk-taking scales ask about specific domains — financial risk, physical risk, and so on — and don't always agree with each other. Da Chi Zhang, Scott Highhouse, and Christopher Nye set out to build a short, general measure of the underlying trait-level willingness to take risks that shows up across domains, publishing GRiPS in the Journal of Behavioral Decision Making in 2018 after testing it across several validation samples.
A single, broad disposition toward risk-taking — not a measure of recklessness, poor judgment, or risk in any one specific domain (financial, physical, etc.).
8 items on a 1–5 scale (strongly disagree → strongly agree), no reverse-scored items. Overall score = mean of all 8 items (1.0–5.0).
Robert Vallerand and colleagues introduced the Dualistic Model of Passion in a 2003 Journal of Personality and Social Psychology paper, distinguishing two ways a genuinely loved activity can fit into someone's life: harmoniously (freely chosen, in balance with everything else) or obsessively (compulsive, hard to step back from, even when it's something the person loves). We use the refined 12-item version most commonly used in later research, plus a 5-item check that confirms the named activity is a real passion and not just a casual interest.
Rated against one activity the student names themselves — the same 17 items would produce different scores for a different activity.
17 items on a 1–7 scale (do not agree at all → completely agree). Harmonious and Obsessive Passion are each the mean of 6 items; Passion Criteria is the mean of the remaining 5. The Passion Scale © Robert J. Vallerand, 2003.
Philip Podsakoff and colleagues published the Transformational Leadership Inventory (TLI) in Leadership Quarterly in 1990, breaking transformational leadership down into six concrete behaviors. It's normally a 360°-style tool — followers rate their supervisor — but since Lens only has a student's own perspective, we adapted all 23 items into first-person self-report form ("I..." instead of "My supervisor..."), preserving each item's original meaning. We chose the TLI over commercially-licensed alternatives like the MLQ specifically because it's a peer-reviewed academic scale, not a paid proprietary instrument.
23 items on a 1–7 scale (strongly disagree → strongly agree). Two items in Individualized Support are reverse-scored. Each of the six subscales is the mean of its own items; an overall score is the mean of all six subscale means. Treat results as a self-perceived leadership style, not a validated 360° assessment.
Not a psychometric assessment — real, self-reported, unverified data, kept as its own distinct step rather than a card in the assessment grid above. Standardized test scores are a well-established, independently validated proxy for quantitative and verbal reasoning; GPA adds real-world context on effort, consistency, and school environment alongside the psychometric instruments. SAT/ACT Math currently carries the fluid-reasoning (Gf) signal in your profile, since Matrix Reasoning (below) is on hold pending licensing.
These three cognitive tasks are drawn from the Cattell-Horn-Carroll (CHC) model of cognitive abilities — the leading psychometric framework for understanding intelligence — but none of them are currently part of the required battery. Matrix Reasoning is on hold pending licensing of the official item set (see below). 3D Rotation and the Remote Associates Test were retired from the required battery entirely: with GRiPS, Passion, and Leadership Styles added, the personality/values/motivation side of the profile carries the synthesis on its own, and both tasks would need real validated item banks to be more than a stand-in. We keep them documented here, and the pages themselves stay live, in case we revisit fluid/spatial/associative cognitive measurement as a more fully-built-out direction later.
David Condon and William Revelle at the University of Chicago created the International Cognitive Ability Resource (ICAR) in 2014 to provide freely available, psychometrically validated cognitive tests as alternatives to expensive proprietary instruments (e.g. Raven's Progressive Matrices, WAIS). The matrix reasoning items directly parallel Raven's in format and measure the same construct — fluid intelligence (Gf): the ability to reason through novel problems without relying on prior knowledge. Gf is the single best predictor of academic and occupational performance among all cognitive abilities.
Each item presents a 3×3 grid of patterns with one cell missing. The respondent selects the correct completion from six options. Success requires inductive reasoning, identifying abstract rules, and applying them.
Number correct out of 16 items. No time limit. Unlike self-report measures, there are objectively right and wrong answers.
This assessment is temporarily removed from the required battery. The puzzles here are our own items, built in the style of ICAR-format matrix reasoning — not the official, validated ICAR item bank. We've reached out to the ICAR project's current maintainers (Cambridge Psychometrics Centre and TU Dortmund) to ask about licensing the real item set, and we're holding this test out of the required 9 until we hear back — using our own uncalibrated items to make a claim about fluid reasoning felt like exactly the kind of overclaiming we don't want to do. In the meantime, SAT/ACT Math (collected as part of Academic Data) stands in as the fluid/quantitative reasoning signal in your profile — it's a well-established, independently validated proxy. For general context, on the actual published ICAR Matrix Reasoning items, a large sample (17,000–35,000 respondents per item) answered correctly on about 52% of items on average, ranging from 28% to 77% depending on the item's difficulty (Condon & Revelle, 2014, icar-project.org).
Mental rotation tasks were pioneered by Roger Shepard and Jacqueline Metzler (1971), who discovered that the time to determine if two 3D objects are the same increases linearly with the angular disparity between them — suggesting the mind literally "rotates" a mental image. Sandra Vandenberg and Allan Kuse formalised a paper-based test in 1978. The ICAR version (2014) adapts this for computer administration. Spatial ability is one of the most job-relevant cognitive traits, particularly in STEM, engineering, surgery, and design.
The ability to mentally rotate 3D objects and determine if two presented views show the same object or mirror images. Measures visual-spatial processing (Gv) — a broad ability that is partially independent of verbal and fluid intelligence.
Number correct. Each item shows a target object and several alternatives; respondents identify which alternatives match the target (some items have multiple correct answers, scored conservatively).
This assessment has been retired from the required battery entirely — it's no longer part of what a student completes. The shapes in Lens's 3D Rotation assessment were our own procedurally-generated cube figures, built in the style of the ICAR spatial battery and the classic Vandenberg & Kuse task — they were not the official, validated item bank, and hadn't been calibrated against real respondents, so no population norm was ever applied to scores. The page stays live, unlinked, in case we revisit spatial measurement with a real validated item bank later.
Sarnoff Mednick introduced the RAT in 1962 based on his associative theory of creativity: creative people have flatter associative hierarchies — meaning they can access a broader, more remote set of associations for any given concept. The RAT operationalises this by presenting problems that require finding a single word connecting three seemingly unrelated words (e.g., pine · crab · sauce → apple). It has become one of the most widely used measures of convergent creative thinking in research. Unlike divergent creativity tests, there is a single correct answer for each item.
Associative fluency (Glr) — the ability to retrieve and connect distantly related concepts stored in long-term memory. Often described as measuring convergent creativity: creative thinking that converges on a single elegant solution rather than generating many options.
Number correct within the time limit. Each correct answer requires finding the one word that forms a compound word or common phrase with each of the three cue words. Difficulty varies across items; harder items require more remote associations.
This assessment has been retired from the required battery entirely — it's no longer part of what a student completes. The word triads in Lens's Remote Associates Test were genuine, published items — Compound Remote Associate problems from Bowden, E.M., & Jung-Beeman, M. (2003), "Normative data for 144 compound remote associate problems," Behavior Research Methods, Instruments, & Computers, 35(4), 634–639. That study tested all 144 items on real participants and published exactly how many solved each one; we used a 30-item subset chosen to span that reported difficulty range. Because our subset and timing/format were our own selection rather than a full re-administration of the validated study, no formal population norm was ever applied to scores here. The page stays live, unlinked, in case we revisit associative-fluency measurement later.
Your profile isn't a free-associated impression of your answers. It follows a specific, research-backed sequence: score everything independently first, look for where the evidence converges, and only then write the narrative.
Each of the 9 assessments (plus your academic data) is scored using its own published scoring method — nothing is blended or averaged across instruments at this stage. This keeps each measurement clean and independently checkable, the same way a lab keeps individual test results separate before a doctor reads them together.
The strongest, most trustworthy statements in your profile are the ones where two or more unrelated instruments point the same direction — for example, a personality trait and a cognitive score reinforcing each other. A single number from a single test is weak evidence on its own; three independent measures agreeing is strong evidence. This "structure first, synthesis second" sequence mirrors a well-documented finding in decision research: judgment that follows structured, independently-scored dimensions consistently outperforms holistic, first-impression judgment — including in high-stakes settings like military personnel selection.
Source: Kahneman, D. — Thinking, Fast and Slow (2011); Kahneman, Sibony & Sunstein — Noise: A Flaw in Human Judgment (2021).
Once your scores are computed and the convergent patterns are identified, we use Claude (built by Anthropic) to turn that structured evidence into a warm, readable narrative — not to re-analyze your raw answers from scratch. The model is given an explicit framework to follow (what counts as a strong signal, how to handle scores that are still developing at your age, what tone to use) rather than being asked for an open-ended opinion. Calibrated AI judgment, applied consistently, has been shown to match trained human raters closely on structured evaluation tasks — the same property that makes the structure-first approach work in the first place.
What this isn't: your profile is intended for personal reflection and as a starting point for conversation with a parent, teacher, or advisor. It is not a diagnostic tool, not a clinical assessment, and not a prediction of who you'll become. Every instrument in the battery was designed to describe patterns in large populations — applying any single result to one person always carries uncertainty, which is exactly why the synthesis looks for convergence across multiple independent measures rather than leaning on any one score.
Once your profile is generated, you can optionally get general guidance on the types of colleges and majors that tend to fit people like you — never specific institutions. Here's the published research that guidance is built on.
In a study tracking students through college, Big Five personality traits accounted for 45% of the variation in life satisfaction, with additional contributions from sense of identity and college satisfaction specifically. Who students are when they arrive on campus appears to meaningfully shape how satisfied they end up being — which is why personality, not just interests, belongs in a college-fit conversation.
Source: Research in Higher Education (2004).
A 2024 study of 4,753 first-year students across 12 North American colleges — large public universities and small private colleges alike — found that more extraverted, less neurotic, and less open students were more likely to attend (and feel they belonged at) large schools. Extraversion and agreeableness predicted belonging more strongly, and that relationship was even stronger at large schools specifically. The practical read: students who build belonging through a few close connections rather than broad social integration are structurally better supported at smaller institutions, where that kind of connection is easier to find.
Source: Stubblebine, Gopalan & Brady (2024), PLoS ONE 19(1): e0295436 — open access.
A study of 8,475 first-year students at 46 institutions found Openness positively associated with peer interaction, course engagement, and second-year retention. Openness seems to matter more for how well a student adjusts and integrates than for raw grades — meaning the institutional question isn't just "how rigorous," but whether the curriculum structure gives an open, curious student room to follow that curiosity.
Source: Bowman, N.A. (2014), Innovative Higher Education.
Across multiple studies, including a cross-classified analysis of over 30,000 students, an institution's support for student autonomy — real choice and voice in coursework — predicts stronger intrinsic motivation and fewer depressive symptoms. A student whose autonomy needs are currently frustrated is a strong candidate for an environment that offers more flexibility and choice, not less.
Source: Self-Determination Theory research; Frontiers in Psychology; ScienceDirect cross-classified path analysis.
This is the finding that most shapes how we frame college guidance: a review synthesizing over 1,800 peer-reviewed studies found little evidence that selectivity relates to self-reported learning gains — a result that has held for 40–50 years. Separately, Gallup-Purdue found no relationship between selectivity and either workplace engagement or general wellbeing after graduation (with one exception: private for-profit institutions, which score lower on both). Students admitted to more selective schools who chose less selective ones did just as well financially. This is why we describe types of environments rather than tiers of prestige — it's what the evidence actually supports.
Source: Challenge Success, Stanford Graduate School of Education (2018), synthesizing Mayhew et al. (2016); Gallup-Purdue Index; Dale & Krueger.
The Gallup-Purdue Index identified six college experiences that predict both workplace engagement and life satisfaction after graduation: a professor who makes learning exciting, professors who care about students personally, a mentor who encourages personal goals, a project spanning multiple semesters, a relevant internship, and active extracurricular involvement. Only 3% of graduates experience all six. Institution size and faculty-to-student ratio meaningfully affect the odds of experiencing the first three — which is a more useful lens than a school's ranking.
Source: Gallup-Purdue Index, cited in the Challenge Success white paper (2018).
Students with more interdependence-oriented values (family, tradition, community) can experience a real mismatch at institutions built around highly individualistic norms of independence — a pattern documented in cultural mismatch research, particularly among first-generation students. This isn't about one set of values being better; it's about whether a campus's culture matches what already feels natural to a given student.
Source: Cultural Mismatch Theory, referenced in the Challenge Success white paper and broader admissions literature.
Holland's person-environment fit theory — the basis for the RIASEC instrument in your battery — has decades of evidence that matching a person's interest type to their field of study and work predicts satisfaction, stability, and achievement. It's the anchor for any major recommendation, with everything else adding nuance on top.
Source: Holland, J. (1959–1997).
The Theory of Work Adjustment holds that good fit requires two matches: a person's abilities meeting what a field actually demands, and a field meeting a person's needs. Reasoning ability points toward different kinds of fields — quantitative/analytical versus more open, interdisciplinary ones — independent of stated interests. The U.S. Department of Labor's O*NET database catalogs the ability requirements of essentially every occupation, giving this kind of matching a real, published evidence base. In Lens's current battery, this layer is carried by SAT/ACT Math (a Gf proxy) alone, now that 3D Rotation and the Remote Associates Test have been retired — see "Currently On Hold / Retired" above.
Source: Dawis & Lofquist — Theory of Work Adjustment; O*NET OnLine, U.S. Department of Labor.
A meta-analysis spanning cumulative samples of over 70,000 students found Conscientiousness to be the strongest Big Five predictor of academic performance — and, importantly, it predicts performance independently of cognitive ability, not as a proxy for it. It's less a major-selecting trait and more a general asset worth being aware of, whatever direction a student chooses.
Source: Poropat, A.E. (2009), Psychological Bulletin.
Social Cognitive Career Theory finds that self-efficacy is a causal precursor to stated interests: what people believe they're good at shapes what they say they're interested in, sometimes before their actual ability is factored in. A student with strong reasoning ability but lower self-reported confidence may be underselling their own fit for a more demanding field — worth naming gently, not as a correction.
Source: Lent, Brown & Hackett — Social Cognitive Career Theory.
Full citations for everything referenced on this page.