est trials at duke

J.B. Rhine Ran 90,000 ESP Trials at Duke. After 1934, Only He Could Reliably Reproduce the Results

19 Min Read

J.B. Rhine did not begin as a paranormal researcher. He arrived at Duke University with a doctorate in botany and an academic career that, on paper, had very little to do with telepathy, clairvoyance, or precognition. What changed the direction of his life was a 1922 lecture by Arthur Conan Doyle, delivered on the same tour in which Doyle was promoting spiritualism and arguing publicly that the evidence for communication with the dead deserved serious consideration.

Rhine attended one of the lectures. The experience did not persuade him that spiritualism was true, but it did persuade him that the underlying question might be worth investigating scientifically. That distinction became the defining feature of his career. Instead of beginning with mediums and séances, Rhine wanted to know whether alleged psychic abilities could be isolated, measured, repeated, and subjected to statistical analysis.

est trials at duke 1

The possibility became institutional when William McDougall recruited him. McDougall was not simply another paranormal enthusiast, he was a Harvard-trained psychologist and a former president of the Society for Psychical Research in Britain, brought to Duke specifically to build its new psychology department, and he sought Rhine out for exactly this purpose. In 1930, with McDougall’s backing, Rhine founded the Parapsychology Laboratory. The methods were deliberately mundane. There were cards. There were guesses. There were controlled trials. There were numbers.

- Signal Intercept -
est trials at duke 2

The extraordinary claim was that the numbers were not behaving as chance predicted. And that is where the story becomes difficult, because Rhine did something genuinely important. He helped turn psychical research from an assortment of séances, mediums, anecdotal experiences, and spontaneous reports into a research program that could at least be argued about statistically. But the same experiments that made ESP look experimentally approachable also produced the methodological problem that would follow the field for decades: when the effect appeared, could researchers actually establish that ESP was responsible for producing it? That question never received the clean answer Rhine wanted.

From Testimony to a Testable Question

Nineteenth-century spiritualism had generated enormous quantities of testimony. Mediums claimed to communicate with the dead. Sitters reported voices, apparitions, automatic writing, and information apparently unavailable through ordinary means. Investigators attempted to catch fraud, and some did. Others remained convinced that at least some phenomena couldn’t be explained by deception. But testimony has a fundamental limitation: even when a witness is honest, the event being described is already filtered through memory, expectation, interpretation, and circumstance.

est trials at duke 3

Rhine’s question was narrower and, scientifically, considerably better formed. Could an alleged paranormal ability produce a measurable deviation from chance under controlled conditions? That framing creates the possibility of failure, if the phenomenon doesn’t occur under the specified conditions, the experiment has at least the potential to count against the hypothesis, which no séance report ever really could.

The tool that made this possible was designed with Duke perceptual psychologist Karl Zener, and it’s worth being specific about why the cards mattered rather than just naming them. A deck of 25 cards, five each of five symbols, a circle, a square, a cross, a set of wavy lines, and a star, gave Rhine something a séance never offered: a fixed, known chance baseline. With five possible symbols, random guessing predicted 20 percent accuracy. That turned a vague, unfalsifiable question, “does this person have psychic ability?”, into a specific, checkable one: does this person’s guessing rate exceed 20 percent by more than chance would produce, and by how much? That reframing, from an amorphous claim into a repeatable experimental protocol with a known baseline, is Rhine’s real methodological contribution, independent of what the results eventually showed.

est trials at duke 4

If someone consistently performed well above that baseline, several explanations remained available: ordinary guessing luck, flaws in procedure, sensory leakage, statistical artifact, selective reporting, experimenter influence, or something genuinely anomalous. ESP was only one possibility among several, and Rhine’s experiments are sometimes remembered as though the experiment itself contained the conclusion. It didn’t. The experiment established a statistical question. The interpretation of the deviation was, and remained, the difficult part.

When the Numbers Started Looking Strange

The early testing was extensive. In 1931 alone, Rhine ran roughly 10,000 individual card trials across 63 Duke students, most of whom scored at or near chance. A few didn’t. The name that built the lab’s early internal reputation was A.J. Linzmayer, a sophomore whose scores ran persistently above what chance predicted across repeated sessions. The result that made the lab nationally famous came in 1933 and 1934, when Rhine’s research assistant J.G. Pratt tested divinity student Hubert Pearce across four separate series, with the target cards shuffled and their order recorded in a locked room 100 to 250 yards from where Pearce sat guessing, a physical separation specifically designed to rule out any ordinary sensory channel between the cards and the guesser. Pearce’s overall accuracy across the Pearce-Pratt series came out to roughly 30 percent against the 20 percent chance predicted, a gap that, given the number of trials involved, wasn’t easily dismissed as ordinary statistical noise.

- Signal Intercept -
est trials at duke 5

Rhine published the first edition of Extra-Sensory Perception in 1934, introducing that term and popularizing the word “parapsychology,” which the German researcher Max Dessoir had actually coined decades earlier. Historians of the field generally treat the book as the starting point of modern experimental parapsychology, the moment the subject moved from anecdote into a university laboratory with statistical controls. His 1937 popular book, New Frontiers of the Mind, brought the lab’s findings to a general readership.

est trials at duke 6

But the important word throughout all of this is statistically. A result can be statistically significant without its underlying cause being established. Statistical significance asks whether an observed outcome would be unusual under the study’s specified chance model. It doesn’t, by itself, identify the mechanism responsible for the deviation. If a subject gets 30 percent rather than 20 percent correct, that’s interesting. If the same pattern appears repeatedly under controlled conditions, it becomes considerably more interesting. But if the result disappears when a different experimenter conducts the test, or when procedures change slightly, the question changes entirely. The issue is no longer simply whether the original numbers were unusual. It’s whether the experimental system itself was generating the effect.

Replication Was Supposed to Settle It

In science, an extraordinary result becomes considerably more persuasive when independent researchers can reproduce it, and Rhine’s work didn’t stand alone. In the five years following his first publication, 33 independent replication attempts were conducted at other universities in the US and Europe. A later statistical review reported that 20 of those 33, roughly 61 percent, produced statistically significant results, against the 5 percent chance alone would predict. That figure became important within parapsychology precisely because it appeared to show Rhine’s findings weren’t a single laboratory’s anomaly.

est trials at duke 7

But the number requires careful handling. A statistically significant result isn’t synonymous with a successful replication of the underlying phenomenon. The precise experimental conditions, statistical assumptions, and methodological quality of each replication all matter, and when a field contains many experiments, some apparently significant results can occur by chance even when the underlying phenomenon doesn’t exist, particularly if researchers are running many analyses and selectively emphasizing the successful ones. The replication record therefore didn’t produce the simple conclusion that ESP had been reproduced twenty times. It produced something more complicated: repeated, statistically unusual findings, and an open question about whether they reflected one underlying anomalous phenomenon or a collection of methodological effects. That distinction is the dividing line between an interesting result and an established one, and it’s exactly where the deeper problem starts.

The Problem Inside the Laboratory

The most damaging criticism of Rhine’s work never required anyone to accuse him of fraud. It was methodological, and it emerged from within the lab’s own record: outcomes appeared to depend too heavily on who was conducting them. Rhine himself sometimes obtained results that other experimenters, working with the same protocols, struggled to reproduce. Historians of the lab don’t dispute this. After 1934, Rhine found he could reliably reproduce statistically significant card-guessing results in sessions he personally ran, while increasingly, other experimenters, including people within his own lab, couldn’t.

That pattern is a problem for almost any experimental claim. Suppose an effect is genuinely produced by telepathy. One would expect the identity of the experimenter to matter only insofar as the experimental conditions actually differed in some relevant way. Suppose instead the effect is being produced by subtle differences in procedure, expectation, communication, recording, or subject selection, then experimenter identity becomes enormously important, and there’s an even more uncomfortable possibility underneath that: experimenters can influence outcomes without realizing they’re doing so at all. This isn’t unique to parapsychology, human beings are exceptionally good at transmitting information unintentionally, and the stronger the claim, the more carefully that possibility has to be ruled out. The problem wasn’t that Rhine’s numbers were necessarily fraudulent. The problem was that the experimental design had never demonstrated the numbers could only have been produced by ESP. That’s a much higher bar, and cheating, sensory leakage, and inadequate shuffling were all independently identified as plausible alternative explanations wherever a given session’s controls were less than airtight.

est trials at duke 8

A paranormal hypothesis can also accommodate almost any outcome, which is precisely why that flexibility is dangerous rather than reassuring. If success demonstrates ESP, failure can be explained by fatigue, anxiety, a bad relationship between subject and experimenter, or the experimenter’s own skepticism interfering with the result. Once every possible outcome can be absorbed into the theory, the theory becomes very difficult to falsify. This is the same trap this library has flagged in other contexts throughout its own investigations, and Rhine’s program, for all its real methodological seriousness, was not immune to it.

- Signal Intercept -

Not every prominent figure in psychical research accepted Rhine’s framing, either. Eileen Garrett, a well-known medium tested at Duke with Zener cards in 1933, performed at essentially chance level, and rather than accepting that as a null result, she pushed back on the test’s premise itself, arguing that whatever ability she believed she had didn’t operate the way card-guessing assumed it would. That objection deserves to be taken seriously rather than dismissed outright, a badly designed experiment really can fail to measure what it claims to, and experimental validity is a real concern, not a rhetorical dodge. But there’s a corresponding danger on the other side: if every negative result can be waved away by saying the experiment was fundamentally incapable of detecting the phenomenon, the claim becomes increasingly insulated from ever being disconfirmed.

est trials at duke 9

The same standard has to run in both directions. A positive result can’t be accepted as evidence for the paranormal just because investigators believe ordinary explanations have been ruled out, and a negative result can’t be dismissed just because the test supposedly measured the wrong thing. Rhine’s work brought that standard into view. It never fully satisfied it.

What the Numbers Actually Established

The history is often distorted in both directions. One version says Rhine proved ESP. The other says the whole enterprise was meaningless because parapsychology isn’t legitimate science. Neither accurately describes what happened. Rhine demonstrated that it was possible to construct controlled experiments around claims of extrasensory perception and obtain datasets containing statistically unusual deviations from chance. Other researchers attempted replications, and a real share of those attempts also produced unusual statistical results. Those are genuine features of the historical record. But the evidence never established that ESP specifically was the cause, and the methodological criticisms weren’t cosmetic objections tacked on by hostile skeptics after the fact. They went directly to the central inferential problem: whether the experimental design could actually distinguish anomalous information transfer from ordinary error, bias, leakage, and experimenter influence. That distinction is everything.

If a coin is tossed one hundred times and lands heads 70 times, the result deserves investigation. It doesn’t establish that the coin possesses a paranormal tendency toward heads. The numbers tell researchers something happened. They do not automatically tell researchers what happened. That distinction is the whole argument, and it’s worth holding onto past this one article, because it’s the same distinction this library tries to apply to every anomaly it examines: a pattern is not, by itself, a cause.

Why Rhine Still Matters

That leaves Rhine in a considerably more interesting historical position than either believer or debunker allows. He didn’t prove ESP. But he did something that permanently changed the argument, taking claims that had previously lived almost entirely in séances, anecdotes, and reports of extraordinary personal experience, and subjecting at least some of them to controlled, quantitative testing. That was a genuine methodological advance, and it created a standard that eventually turned back against the field itself. Once paranormal claims are expressed statistically, they become vulnerable to statistical criticism. Once experiments are standardized, procedural differences can be directly compared. Once replication becomes the criterion, researchers can no longer rely indefinitely on one extraordinary result, including their own.

The 1938 American Psychological Association convention gave the dispute its most formal hearing, a rare instance of a mainstream scientific body devoting a structured session specifically to “Experimental Methods of ESP Research” rather than simply ignoring the claim. Rhine and his statistician sat on one side, alongside the respected psychologist Gardner Murphy, who argued the methodology deserved serious engagement without personally endorsing its conclusions. Three established critics argued the other side. The panel didn’t produce consensus. What it produced was decades of continued, structured disagreement, exactly the kind of ongoing argument a genuinely testable claim generates, and a séance never could.

- Signal Intercept -

The Parapsychology Laboratory operated at Duke for thirty-five years before Rhine’s 1965 retirement, after which the work continued as an independent nonprofit, now the Rhine Research Center, still operating in Durham today, decades after Rhine’s own death in 1980. Duke’s Rubenstein Library holds roughly 700 boxes of the lab’s original records and correspondence, a real, substantial primary-source archive historians continue to work through.

est trials at duke 10

The institutional afterlife matters for the same reason the 1938 debate does: it’s evidence that Rhine built something durable enough to keep arguing about, which is a different, more modest achievement than building something that settled the argument.

The Question Rhine Left Behind

There’s something almost fitting about the fact that Rhine’s greatest legacy isn’t an answer. It’s a better question. Before Rhine, the debate over psychic phenomena could too easily stay a contest between testimony and disbelief: someone experienced something extraordinary, someone else insisted it couldn’t have happened, and neither side had a shared procedure for finding out who was right.

Rhine introduced another possibility. Count it. Repeat it. Control it. Try to make it disappear. Then see what remains. That procedure didn’t produce the definitive proof of ESP Rhine hoped for. But neither did every experiment simply collapse into nothing, and the historical record is more uncomfortable than either conclusion would suggest. Something repeatedly appeared in the numbers. The problem was, and remains, determining what the numbers were actually measuring. Can this be measured. Can it be replicated. Can alternative explanations be eliminated. And then the harder question underneath all three: if an effect survives all of that scrutiny, what exactly does it demonstrate? That is not whether the cards knew the answer. It’s whether the experiment ever demonstrated that anyone, or anything, did.

Share This Article
Leave a Comment