Author's note: A long, occasionally tedious, but very important article about tools for assessing people.

In short — there is no magic bullet.

If not in short — see below.

Warning: there are a lot of words ahead, a few numbers, and slightly fewer illusions.

This is one of the most frequent and most dangerous questions in personnel assessment. It is dangerous not because it should not be asked, but because behind it there almost always lies a hidden hope: perhaps there is, after all, one instrument, one test, one interview, one assessment center, one "magic battery" that will finally make it possible to stop making mistakes about people.

The scientifically honest answer is not a comfortable one: no, a one-hundred-percent accurate instrument for assessing people does not exist, and the problem lies not only in the quality of the methods themselves, but also in the nature of what we are trying to measure.

Work is complex, human behavior changes with context, the criteria of success are often imprecise, and any validity in selection is always a probability, not a verdict. This is precisely why professional psychometrics and I/O psychology have long since moved away from the logic of "finding the perfect test" and toward the logic of assembling the most well-grounded combination of instruments for a specific task.

First, it is important to agree on a common language. When researchers write that the validity of a method is, for example, .42 or .28, this does not mean "42% accuracy" or "28% correct decisions."

What is usually meant is criterion-related validity — the correlation between the instrument's result and subsequent job performance.

Even a very good instrument does not "see through a person completely"; it merely provides a more or less strong statistical signal.

For a correlation of .42, the squared correlation is approximately 18% of explained variance — which is already practically useful, but is very far from the fantasy of "complete transparency" of the person.

Furthermore, the success criteria themselves — for example, a manager's ratings — are also imperfect, which means the ceiling of any accuracy is further limited by the quality of how we measure the outcome.

This is precisely what official professional testing standards reflect: validity is not a "property of the test in general," but the justifiability of interpreting results for a specific application, and each use requires its own evidence.

This is why, from a scientific standpoint, the question "which instrument is the most accurate?" must always be reframed: for what role, for what criterion, in what population, at what stage, for what purpose — selection, promotion, development, succession, risk assessment, potential assessment?

One and the same method can behave very differently depending on whether we are assessing a line employee, a future manager, someone at the beginning of their career, or a sitting senior executive.

This is one of the reasons why modern literature devotes so much attention not only to the validity of individual procedures, but also to the design of the selection system as a whole.

What the global research base shows

If we set aside popular HR myths and look at large meta-analyses, the picture is simultaneously sober and interesting.

Older reviews — most notably the work of Schmidt and Hunter — long entrenched in the professional field the idea that general mental ability / cognitive ability tests are the best single predictor of job performance.

In the seminal 1998 paper, an operational validity of .51 was cited for GMA.

But more recent revisions have shown that a number of older estimates were likely inflated due to the methods used to correct for range restriction. In a major revision by Sackett, Zhang, Berry, and Lievens, it was shown that the validities of many methods were overstated by .10–.20, and in their updated matrix the structured interview turned out to be the strongest single predictor among widely used methods — .42 — while for GMA in that same matrix the estimate was .31, for biodata — .38, for integrity tests — .31, for situational judgment tests — .26, and for conscientiousness tests — .19.

Even more striking is a recent meta-analysis by Sackett and colleagues covering the twenty-first century: for GCA and overall job performance, they obtained a corrected validity of .22 — not the "legendary" .51.

This leads to a very important managerial conclusion.

First, there is no longer an unchallenged absolute leader.

Second, the old line of reasoning — "just give a good IQ test and everything will be clear" — sounds far too crude and far too overconfident today.

Cognitive tests remain a serious instrument, but they no longer look like the unquestioned pinnacle they were often made out to be in popular business literature.

Cognitive tests: a strong instrument, but no longer "king of the hill"

I will start here because more myths have accumulated around cognitive tests than around anything else.

Historically, cognitive ability tests genuinely had a very strong reputation. And that reputation did not arise from thin air: these tests are consistently linked to learnability, speed of absorbing new information, solving complex problems, and, in many studies, to overall job performance.

Even in the new meta-analysis, where the estimate is reduced to .22 for the twenty-first century, the authors explicitly state that the link to productivity exists — it is simply smaller in magnitude than previously believed. In Sackett and colleagues' updated integrative matrix, the figure for GMA is .31 — still a strong single predictor, but no longer a dominant one.

The problem with cognitive tests is that organizations often expect too much from them. They do a good job of answering the question of how quickly a person learns and how well they handle tasks requiring analysis and mental processing of information, but they do a much poorer job of describing things like influence style, maturity of judgment, behavioral stability under pressure, the ability to build relationships, manage conflict, tolerate organizational ambiguity, or be a constructive leader.

If a role is rich in precisely these components, relying solely on a cognitive test means taking only one cross-section of the person and pretending that is enough.

Structured interview: possibly the best single instrument among those that are genuinely practical

If someone asked me which single method today looks like the most recognized and practically the most reasonable for most companies, I would answer as follows: a well-designed structured interview is the safest candidate for the role of "best single instrument" — if we are talking not about myths but about the modern evidence base.

In Sackett and colleagues' updated matrix, it is structured interview that comes in first place with a validity of .42.

And an older but still important meta-analysis by McDaniel and colleagues showed that interviews in general "work" considerably better than skeptics long believed, and that structured interviews systematically outperform unstructured ones.

Reviews of this body of research typically cite figures of around .44 for structured interviews versus approximately .33 for unstructured ones, using job performance as the criterion.

Why does this matter so much?

Because a structured interview is not simply "a conversation based on pre-prepared questions" — it is a method that involves job analysis, clear criteria, a consistent logic of questions for all candidates, behavioral anchors or rating guides, a uniform scoring process, and interviewer discipline.

It is precisely the structure that reduces noise, diminishes the influence of liking, charisma, and first impressions, and yields more comparable data.

In other words, practitioners often underestimate not the interview itself, but the price of how structured it actually is.

Work sample tests and job knowledge tests: very strong when the role allows it

If a person already knows how to do the job, the most direct form of assessment is to observe how they perform it, or a realistic fragment of it.

The logic of work sample tests is so natural that organizations often embrace them intuitively — and not without reason.

In more recent reviews, Schmidt and Oh report a validity of .33 for work samples, whereas older estimates were substantially higher, reaching as high as .54; the update by Roth, Bobko, and McFarland reduced the older figure precisely after accumulating broader data.

For tailored job knowledge tests, Schmidt and Oh's review cites a validity of .48.

This is a very good result, but it applies primarily to situations where it makes sense to test existing professional knowledge rather than general potential.

And it is precisely here that the core principle of quality assessment becomes apparent: the strength of an instrument depends on how well it matches the nature of the task. If you are hiring a welder, an accountant, an analyst with a clear set of operational duties, or a manager who needs to solve a real business case, a work sample can provide an incredibly valuable signal.

But if you are trying to use a work sample to predict long-term leadership potential, cultural maturity, or a person's capacity to grow in a different role two years from now, this method ceases to be all-powerful.

It is strong where what needs to be measured is "can they do it" and considerably weaker where what needs to be forecast is "how will they grow, influence others, and behave in a complex, living system."

Situational Judgment Tests: a useful "middle class" of assessment

SJTs are one of the most practical instruments for assessing judgment, social understanding, priorities, and the applied behavioral logic of typical work situations.

In Sackett and colleagues' updated matrix, the validity cited for SJTs is .26 — not the table leader, but a perfectly workable instrument. At the same time, reviews of situational judgment tests (SJTs) specifically emphasize that their strength lies not only in predictive validity (criterion validity), but also in the fact that they frequently deliver incremental validity over and above cognitive ability tests and personality measures.

Moreover, they typically carry a lower risk of systematic selection bias (adverse impact), especially in cases where items are less loaded with cognitive complexity (cognitive loading).

There is one more important nuance: response format (response instructions) matters greatly.

For example:

  • tests focused on knowledge of the "correct" behavior (knowledge-based SJT) are more strongly linked to cognitive ability;
  • whereas tests in which the person chooses how they themselves would act (behavioral tendency SJT) are more strongly linked to personality characteristics such as conscientiousness, agreeableness, and emotional stability.

In practice, this means: SJT rarely serves as the "best single method," but very often turns out to be an excellent second or third element of a battery — especially when an organization wants to assess not only intelligence but also judgment, response to ambiguous situations, and social-behavioral preferences.

Personality tests: useful, but only if we stop expecting magic from them

This is perhaps where the most disappointment accumulates. Because personality questionnaires are very easy to sell as an "X-ray of the person," while research has been speaking in a more modest voice for many years.

The classic meta-analyses by Barrick and Mount (Barrick & Mount), and then by Hurtz and Donovan (Hurtz & Donovan), showed that among the Big Five model, the most consistently stable predictor of job performance is conscientiousness, with typical validity values of around 0.22–0.25 in earlier studies.

More recent higher-order meta-analyses (second-order meta-analysis) also confirm that the relationship between conscientiousness and performance remains at approximately 0.20.

However, there is an important nuance that is often missed. A recent paper by Watrin, Weihrauch, and Wilhelm (Watrin, Weihrauch & Wilhelm) draws attention to the fact that a significant portion of these results come from studies that analyzed employees who were already in the role (concurrent studies with incumbents).

In other words, measurements were taken "here and now" rather than in a predictive logic.

When one attempts to transfer these findings to applicant-based predictive settings, the relationship turns out to be less clear-cut than is commonly assumed.

This is also visible in Sackett and colleagues' updated integrative matrix (Sackett et al.), where the validity of conscientiousness tests is approximately 0.19.

This does not mean that personality tests are useless.

It means that they are especially valuable not as a standalone basis for "yes/no," but as a layer of data about style, tendencies, risks, likely behavioral patterns, and as a complement to other methods.

Where a recruiter or HRD wants a single button to "measure personality and understand everything," that is no longer science — it is an appealing fantasy.

Integrity tests: an underappreciated instrument, especially when risks of destructive behavior matter

To be candid, integrity tests deserve far more attention than they typically receive in business practice.

In the classic meta-analysis by Ones, Viswesvaran, and Schmidt (Ones, Viswesvaran & Schmidt), these tests yielded a mean predictive validity of approximately 0.41 for overall job performance and approximately 0.47 for counterproductive work behavior (CWB) — such as theft, disciplinary violations, absenteeism, and other forms of destructive conduct.

In more recent synthesizing estimates (for example, in Sackett and colleagues' updated matrix, Sackett et al.), a more conservative value of approximately 0.31 is cited — which still makes this instrument one of the strongest single predictors available.

But perhaps the most interesting finding emerges not in isolated figures but in method combinations.

For example:

  • combining a structured interview (SI) with integrity tests yields a validity of approximately 0.53;
  • and combining biodata with a structured interview and integrity tests yields approximately 0.57.

This illustrates very clearly how the principle of modern assessment works: maximum accuracy emerges not from a single "strong" instrument, but from the intelligent combination of methods with different natures of signal.

This is a very important point for companies that hire for roles where the cost of errors is high, where there is access to resources, customer risk, compliance risk, or a high probability of hidden destructive behavior.

If an organization pays no attention whatsoever to integrity-related constructs, it may be missing precisely the layer of risk that blows up only after the offer has been made.

Biodata: a surprisingly strong, but underappreciated method

Biodata is not simply "a questionnaire about the past."

In the professional sense, it is structured biographical information coded in such a way that past experience, stable patterns of choice, achievements, career trajectory, and typical behavioral facts function as predictors of future performance.

In Sackett and colleagues' updated matrix, biodata performs surprisingly well — .38 as a single predictor, which is higher than GMA and integrity in that particular matrix, and second only to the structured interview.

In combinations, biodata also performs well:

  • combining biodata (BD) with a structured interview (SI) yields a validity of approximately 0.52;
  • combining biodata with integrity tests (I) yields approximately 0.44;
  • and combining all three — biodata, structured interview, and integrity tests — yields approximately 0.57.

Why is this method still not widely popular?

Because it requires smart construction and discipline.

Poorly assembled biographical data is simply a long questionnaire.

Well-constructed biodata is one way of turning a person's past into a predictive signal.

But this only works when the organization understands which specific past facts are genuinely relevant to future job performance.

Assessment Center: a strong method, but not all-powerful

Assessment Centers are widely valued for the effect of a "three-dimensional portrait."

And rightly so.

But it is precisely around them that the most romanticism tends to accumulate. In good hands, an Assessment Center is indeed one of the most substantive methods — especially for managers and complex roles.

However, from the standpoint of criterion-related validity, it is not a magical champion.

In a meta-analysis by Hermelin, Lievens, and Robertson, the corrected correlation between overall assessment rating and supervisory performance ratings was .28, with a 95% confidence interval of .24–.32.

An older meta-analysis by Gaugler and colleagues yielded a higher estimate of around .37. The difference itself is useful as a reminder: even for strong assessment methods, final figures depend on the quality of the studies, the design of the method, and the correction approaches used.

Assessment Centers have one more important feature.

They are very powerful not only as predictors, but also as rich behavioral material for decisions about development, succession, and potential.

But this is precisely why companies often commit a logical error: assuming that because an AC is "complex and expensive," it is automatically "more accurate than anything else."

Science does not support that conclusion. A good AC is valuable, but it does not eliminate the need to combine it with other data, nor does it exempt one from questions about criteria, validity, and observer quality.

Unstructured interview, reference checks, and graphology: the zone of greatest self-deception

Here is where the most illusions in the assessment of people are concentrated.

With unstructured interviews, the situation is fairly clear: they are systematically inferior to structured interviews in predictive value.

The problem lies not in the conversational format itself, but in the absence of structure — uniform questions, criteria, and scoring rules.

As a result, the decision begins to depend not so much on the candidate as on the interviewer's perception: first impressions, liking, communication style.

This is why the belief in "I simply have a good read on people" most often turns out to be not a professional strength, but a straightforward overestimation of one's own intuition.

With reference checks, the picture is more nuanced.

Formally, they are a source of additional information, but in real practice they often turn out to be rather weak.

The reason lies in the human factor and the context: former employers are cautious in their wording, avoid direct evaluations, and sometimes simply provide the most neutral feedback possible. In the end, references are useful as a check on facts and risk signals, but very rarely yield deep understanding of the person.

Graphology, however, is an entirely different story.

From a scientific standpoint, its validity in assessing professional effectiveness is not confirmed.

Research shows that graphology specialists do not produce more accurate predictions than people without specialized training when it comes to forecasting job performance.

To put it plainly: if a company tries to assess candidates by handwriting, it is relying not on evidence-based methods, but on a persistent but unconfirmed myth.

Is there one universally recognized "accurate" instrument?

If the answer is to be maximally honest: the single most accurate instrument "in general" does not exist.

But there are methods that work better than others.

Based on modern large-scale reviews and validity revisions, among widely used single instruments, the structured interview comes out on top with a validity of approximately 0.42.

When it comes to the most extensively studied predictor in the history of research, that is unquestionably general cognitive ability.

However, more recent data show that its predictive power is lower than previously believed and requires more cautious interpretation.

If the nature of the role allows for practical tasks, job knowledge tests and work samples can provide a very strong signal — especially when assessing experienced specialists.

If destructive behavior risk and reliability are critical for the role, integrity tests — which demonstrate high practical value in this domain — should be taken into account.

If the goal is to gain a richer understanding of a manager's behavior, an Assessment Center may be a reasonable element of the system — particularly when working with leadership positions.

But for all of that, there is a fundamentally important limitation: none of these instruments, used in isolation, provides grounds for claiming complete accuracy in assessing a person.

The most important conclusion: accuracy lives not in a single instrument, but in combination

And so we arrive at a conclusion that is simultaneously uncomfortable and liberating for business: there is no magic bullet, but there are good systems.

This is, arguably, the central idea of modern evidence-based assessment.

There is no need to look for a single miracle test.

What is needed is to design a battery of methods so that they yield different, not fully overlapping signals.

This is precisely why combinations often outperform single instruments. In Sackett and colleagues' updated matrix, structured interview + integrity yields .53, biodata + structured interview — .52, GMA + structured interview — .48, and biodata + structured interview + integrity — .57.

That is to say, the strength emerges not because we "piled on more things," but because we combined instruments with different natures of signal.

A good assessment system typically answers several different questions simultaneously. For example:

  • is the person capable of learning quickly and making sense of complex things?
  • How do they reason about work situations?
  • What have they actually done before? How do they behave in a live simulation?
  • How pronounced are their reliability and self-regulation? — What behavioral risks are they prone to?
  • How well does their management style match the demands of the role?

No single method covers all of this at once.

But several well-chosen instruments can.

One more honest note: even the best battery does not replace professional judgment

This is important to state separately.

Professional assessment is not a competition between "the machine and the human." Good instruments are not there to remove expert judgment — they are there to make it less blind and less overconfident.

The most dangerous combination in assessment is not the absence of tests. The most dangerous combination is weak methods plus the assessor's strong confidence in their own infallibility.

Science is precisely what is needed to discipline that confidence.

To summarize very briefly, the conclusion is as follows.

No, it is not possible to obtain 100% accurate data about a person using a single instrument.

And even with a set of instruments, it is not possible, literally speaking, to "know everything."

But it is entirely possible to substantially reduce the probability of error — by stopping the search for magic and instead building assessment as a system: tailored to the role, the criteria, the risks, the context, and grounded in methods for which there is a genuinely serious research base.

And if a single most recognized method of today must be named, I would put my money on the structured interview. And if the best principle of assessment in general must be named, it sounds different: not one instrument, but an intelligent combination of several valid instruments.