Liking each other is not evidence.
Kalibr measures two people separately, then computes how they fit. Each person takes one 191-item psychometric assessment, about 20 minutes. A deterministic engine scores each of them across 10 dimensions. Kalibr then compares the two scored profiles and produces a cohesion report.
A cohesion report cannot exist until both people have completed the assessment and unlocked their own results. Both participate, both consent, and both read the same report. There is no version where one person buys a reading of somebody who did not take it.
The comparison is arithmetic
Most compatibility products hand two personality summaries to a language model and print its impression. Kalibr does the opposite. Both people are scored by a fixed algorithm, with no model involved at the measurement stage. The pairwise layer is arithmetic over those scores, with an explicit weight per dimension, explicit thresholds for what counts as a shared gap or a natural cover or visible friction, and named rules that either fire or do not. The same two assessments produce the same numbers every time. A model writes the sentences, and only once the numbers exist.
What the cohesion report contains
- A cohesion score out of 100 with its band: Strong Foundation, Functional Pairing, Mixed Signals, Significant Tensions, or High Risk.
- A pairwise score for each of the 10 dimensions, and whether it is similarity-driven, complementary, or divergence-driven.
- A symmetry check, with an echo-chamber flag when the pair is too alike to challenge each other.
- Risk concentration: whether friction is spread across dimensions or acute in one.
- Top shared strengths, top tensions, shared gaps, natural covers, and hidden risks.
- Friction narratives, working agreements, and a final recommendation.
The 10 dimensions are Philosophy Cohesion, Drive Alignment, Bonding Index, Adaptive Intelligence, Volatility Vector, Ambiguity Tolerance, Influence Style, Feedback Orientation, Temporal Orientation, and Energy Resilience. Grit and resilient coping are supplementary scores that feed several of them; they are not themselves dimensions.
The science, and its limits
The battery is 191 items in one sitting, about twenty minutes. Every dimension traces to a published, independently validated instrument; the scoring engine and the pairwise layer built on them are Kalibr's own.
Three tiers, kept apart on purpose. Established: using validated personality measurement to predict team outcomes is one of the most replicated findings in organisational psychology. Defensible by design: the dimension composites, the per-dimension weights and the pairwise curve follow from the literature and explicit design logic, but have not been independently validated against outcome data. Open hypothesis: no published study has shown that a pairwise score computed before two people have worked together predicts the quality of their working relationship six months in. Construct validity exists here; criterion validity does not yet.
Kalibr measures working-style preferences and interpersonal dynamics. It does not measure intelligence, skill, education or training, and it is not a hiring decision tool. No cohesion score it produces has been shown to predict whether a partnership lasts.
Kalibr scores individuals, not demographic categories. Two meta-analyses on generational differences at work (Costanza et al., 2012, N approximately 20,000; Ravid, Costanza and Romero, 2025) find the differences between generations on job satisfaction, organisational commitment and intent to leave to be small to nil, shrinking further once age is controlled for, with variation inside a generation exceeding variation between generations. Friction attributed to somebody's generation is usually a difference between two particular people, which is the thing this engine measures.
Pricing
Each person unlocks their own results once, at $75, charged as ₹6,300. The first cohesion report between two unlocked people is free; each further report is $50. No subscription. Stated plainly: a pair reading their first cohesion report has paid $75 each and nothing for the report.
Frequently asked questions
Will my co-founder and I actually work well together?
Liking each other predicts very little about whether a working partnership survives its first hard year, and the usual way to find out is to live it. Both people complete the same validated instrument separately, a deterministic algorithm scores each of them, and the fit between the two score sets is computed rather than described. Kalibr does not tell you whether to go ahead, and no cohesion score has been shown to predict whether a partnership lasts.
Can I get a compatibility report about someone without them taking the assessment?
No. A cohesion report compares two sets of scores, so it cannot exist until both people have completed the assessment and unlocked their own results. This is enforced in the product: the pairing cannot be created until the other person holds a completed scored profile.
Is a computed compatibility score better than an AI's opinion of two personality summaries?
It is a different kind of claim. A model's impression of two summaries can change between runs and cannot be audited. A computed comparison is arithmetic with explicit weights, explicit thresholds and named rules, so the same two assessments produce the same numbers every time. Neither approach has been validated against long-term outcome data.
Why do teams take months to start working well together?
Tuckman's model describes a sequence groups pass through: forming, storming, norming, performing. The storming phase is where differences that were invisible at the start begin to cost something, and most pairs discover their mismatch by hitting it. A pairwise assessment does not shorten the sequence, and Kalibr makes no claim that it does. What it can do is name, in advance, which dimension the storming will be about.
Is psychometrics scientifically valid?
It depends on the instrument. MBTI has weak validity. The Big Five and its public-domain implementations such as the IPIP have held up across six decades of replication, with validity coefficients in the 0.3 to 0.5 range for predicting relevant outcomes. Kalibr draws on four published instruments and separates publicly between what is established, what is defensible by design, and what is an open hypothesis.
What people said about the AI constitution
Unedited responses from people who completed the assessment and loaded the constitution it generates. They are about the constitution, which is a secondary feature, and not about the cohesion report, which is what this site sells. Roles are given as stated; no outcome was measured for any of these people, and none of them is a case study.
"It challenges me and pushes back at the right moments, rightly identifies when I spiral into analysis paralysis, and is direct when needed. Combined with my project-level CLAUDE.md instructions it's turning out to be really good."
"I didn't realise how much of an echo chamber Gemini had become until I loaded my Kalibr constitution. It flagged my documented tendency for impulsive scope creep and demanded a 30-day risk mitigation plan. It literally feels like having a skeptical board member in my terminal."
"I've taken MBTI, DISC, and CliftonStrengths. All gave me 30 pages of corporate fluff about how 'visionary' I am. Kalibr gave me one page that said my extreme tolerance for ambiguity is actively creating operational chaos for my direct reports. Brutal, but it's the first assessment I've ever actually used to change how I manage."
"'Decisions that require straightforward execution may be over-theorised.' This has actually happened with me and one of my employees twice. My manager had to step in both times. Impressive your system found this."
Privacy
Your responses and scores are yours. A cohesion report is visible only to the two people in it, and nobody obtains a reading of a person who did not take the assessment themselves. Kalibr staff hold an administrative surface used to regenerate a stored report; in an organisational deployment, the inviting organisation may hold the reports of people it invited, as set out in the privacy policy.
References
Barrick, M. R. and Mount, M. K. (1991). The Big Five personality dimensions and job performance: a meta-analysis. Personnel Psychology, 44, 1 to 26.
Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: a meta-analysis. Journal of Applied Psychology, 92(3), 595 to 615.
Smith, B. W. et al. (2008). The Brief Resilience Scale: assessing the ability to bounce back. International Journal of Behavioral Medicine, 15, 194 to 200.
Duckworth, A. L. and Quinn, P. D. (2009). Development and validation of the Short Grit Scale (Grit-S). Journal of Personality Assessment.
Goldberg, L. R. et al. (2006). The International Personality Item Pool and the future of public-domain personality measures. Journal of Research in Personality.
Peeters, M. A. G. et al. (2006). Personality and team performance: a meta-analysis. European Journal of Personality.
Dryer, D. C. and Horowitz, L. M. (1997). When do opposites attract? Interpersonal complementarity versus similarity. Journal of Personality and Social Psychology.
Tuckman, B. W. (1965). Developmental sequence in small groups. Psychological Bulletin.
Blog
Writing on psychometrics, AI, and teams.
MBTI is astrology. Here's what actually holds up.
June 2026 · 5 min read
If you've taken the Myers-Briggs test, you were told you're an INTJ or an ENFP and handed a description that felt uncannily accurate. That feeling is the Barnum effect: vague, flattering descriptions that almost anyone will accept as their own. Horoscopes work the same way.
MBTI has a retest reliability problem. Take it twice, six weeks apart, and roughly 50% of people get a different type. That's no better than chance. DISC is marginally better on reliability but still predicts almost nothing about how someone behaves under pressure.
The Big Five has been replicated across six decades of research, across cultures, in laboratory and field settings. The five factors predict job performance, relationship stability, health outcomes, and behaviour under stress with validity coefficients in the 0.3 to 0.5 range. A 1991 meta-analysis of 117 validity studies (N≈24,000) found conscientiousness predicts job performance across every occupational group tested.
These aren't destiny. They're base rates for your own behaviour. The point isn't to label you. It's to give you a map specific enough to be actionable: not "you're a planner" but "you over-weight precision and under-weight speed when the cost of delay is high."
Kalibr's output covers 10 dimensions: Philosophy Cohesion, Drive Alignment, Bonding Index, Adaptive Intelligence, Volatility Vector, Ambiguity Tolerance, Influence Style, Feedback Orientation, Temporal Orientation, and Energy Resilience. Scoring is deterministic: identical inputs always produce identical outputs.
Your AI doesn't know you. That's the whole problem.
June 2026 · 6 min read
Everyone's noticed that AI advice is generic. The usual explanation is sycophancy: models are trained with RLHF, which rewards responses that feel good, so the model learns to validate rather than challenge.
That's true, but it misses the more important problem. A model that knew you well could still be sycophantic. The bigger issue is that the model has no information about you at all. It defaults to a median-human prior, which is wrong for any specific person in proportion to how different they are from average.
When you ask a model to evaluate your business idea, it doesn't know you have a documented pattern of over-weighting revenue potential and under-weighting operational complexity. Without that context, it produces answers designed for a generic professional.
Why your team is still storming six months in.
June 2026 · 6 min read
In 1965, psychologist Bruce Tuckman described four stages every new team moves through: forming, storming, norming, performing. Most teams spend far too long in storming. Six months is common for early-stage teams.
The reason isn't bad hiring. It's that the differences driving friction are invisible. When two people disagree repeatedly about how much information they need before making a decision, they don't experience it as a measurable difference in risk tolerance. They experience it as one person being reckless and the other being indecisive.
Some dimensions are weighted more heavily than others when the engine scores a pair, because a gap on them is more consequential: they govern how a pair handles conflict under pressure, disagreement on irreversible decisions, competition for direction, and relational asymmetry. Which they are, and by how much, is part of the scoring model and is not published. Two people far apart on one of them will produce conflict in exactly that domain, repeatedly, until they have an explicit shared model of why. This weighting is a design decision grounded in the constructs, not a validated ranking against outcome data.
The cohesion report gives both people that map before the friction starts. The working agreement doesn't resolve the differences. It names them. Whether that makes the storming phase shorter is not something Kalibr has measured, and it does not claim it.
Visit kalibriq.com to take the assessment. JavaScript is required for the app itself.