McNemar's test

In statistics, McNemar's test is a statistical test used on paired nominal data. It is applied to 2 × 2 contingency tables with a dichotomous trait, with matched pairs of subjects, to determine whether the row and column marginal frequencies are equal. It is named after Quinn McNemar, who introduced it in 1947.
An application of the test in genetics is the transmission disequilibrium test for detecting linkage disequilibrium. The commonly used parameters to assess a diagnostic test in medical sciences are sensitivity and specificity. Sensitivity is the ability of a test to correctly identify the people with disease. Specificity is the ability of the test to correctly identify those without the disease. Now presume two tests are performed on the same group of patients. And also presume that these tests have identical sensitivity and specificity. In this situation one is carried away by these findings and presume that both the tests are equivalent. However this may not be the case. For this we have to study the patients with disease and patients without disease. We also have to find out where these two tests disagree with each other. This is precisely the basis of McNemar's test. This test compares the sensitivity and specificity of two diagnostic tests on the same group of patients.

Definition

The test is applied to a 2 × 2 contingency table, which tabulates the outcomes of two tests on a sample of n subjects, as follows.

	Test 2 positive	Test 2 negative	Row total
Test 1 positive	a	b	a + b
Test 1 negative	c	d	c + d
Column total	a + c	b + d	n

The null hypothesis of marginal homogeneity states that the two marginal probabilities for each outcome are the same, i.e. p_a + p_b = p_a + p_c and p_c + p_d = p_b + p_d.
Thus the null and alternative hypotheses are
Here p_a, etc., denote the theoretical probability of occurrences in cells with the corresponding label.
The McNemar test statistic is:
Under the null hypothesis, with a sufficiently large number of discordants, has a chi-squared distribution with 1 degree of freedom. If the result is significant, this provides sufficient evidence to reject the null hypothesis, in favour of the alternative hypothesis that p_b ≠ p_c, which would mean that the marginal proportions are significantly different from each other.

Variations

If either b or c is small then is not well-approximated by the chi-squared distribution. An exact binomial test can then be used, where b is compared to a binomial distribution with size parameter n = b + c and p = 0.5. Effectively, the exact binomial test evaluates the imbalance in the discordants b and c. To achieve a two-sided P-value, the P-value of the extreme tail should be multiplied by 2. For b ≥ c:
which is simply twice the binomial distribution cumulative distribution function with p = 0.5 and n = b + c.
Edwards proposed the following continuity corrected version of the McNemar test to approximate the binomial exact-P-value:
The mid-P McNemar test is calculated by subtracting half the probability of the observed b from the exact
one-sided P-value, then double it to obtain the two-sided mid-P-value:
This is equivalent to:
where the second term is the binomial distribution probability mass function and n = b + c. Binomial distribution functions are readily available in common software packages and the McNemar mid-P test can easily be calculated.
The traditional advice has been to use the exact binomial test when b + c < 25. However, simulations have shown both the exact binomial test and the McNemar test with continuity correction to be overly conservative. When b + c < 6, the exact-P-value always exceeds the common significance level 0.05. The original McNemar test was most powerful, but often slightly liberal. The mid-P version was almost as powerful as the asymptotic McNemar test and was not found to exceed the nominal significance level.

Examples

In the first example, a researcher attempts to determine if a drug has an effect on a particular disease. Counts of individuals are given in the table, with the diagnosis before treatment given in the rows, and the diagnosis after treatment in the columns. The test requires the same subjects to be included in the before-and-after measurements.

	After: present	After: absent	Row total
Before: present	101	121	222
Before: absent	59	33	92
Column total	160	154	314

In this example, the null hypothesis of "marginal homogeneity" would mean there was no effect of the treatment. From the above data, the McNemar test statistic:
has the value 21.35, which is extremely unlikely to form the distribution implied by the null hypothesis. Thus the test provides strong evidence to reject the null hypothesis of no treatment effect.
A second example illustrates differences between the asymptotic McNemar test and alternatives. The data table is formatted as before, with different numbers in the cells:

	After: present	After: absent	Row total
Before: present	59	6	65
Before: absent	16	80	96
Column total	75	86	161

With these data, the sample size is not small, however results from the McNemar test and other versions are different. The exact binomial test gives P = 0.053 and McNemar's test with continuity correction gives = 3.68 and P = 0.055. The asymptotic McNemar's test gives = 4.55 and P = 0.033 and the mid-P McNemar's test gives P = 0.035. Both the McNemar's test and mid-P version provide stronger evidence for a statistically significant treatment effect in this second example.

Discussion

An interesting observation when interpreting McNemar's test is that the elements of the main diagonal do not contribute to the decision about whether pre- or post-treatment condition is more favourable. Thus, the sum b + c can be small and statistical power of the tests described above can be low even though the number of pairs a + b + c + d is large.
An extension of McNemar's test exists in situations where independence does not necessarily hold between the pairs; instead, there are clusters of paired data where the pairs in a cluster may not be independent, but independence holds between different clusters. An example is analyzing the effectiveness of a dental procedure; in this case, a pair corresponds to the treatment of an individual tooth in patients who might have multiple teeth treated; the effectiveness of treatment of two teeth in the same patient is not likely to be independent, but the treatment of two teeth in different patients is more likely to be independent.

Information in the pairings

In the 1970s, it was conjectured that retaining one's tonsils might protect against Hodgkin's lymphoma. John Rice wrote:

85 Hodgkin's patients had a sibling of the same sex
who was free of the disease and whose age was within 5 years of
the patient's. These investigators presented the following table:
They calculated a chi-squared statistic had made an error in their analysis by ignoring the pairings. samples were not independent, because the siblings were paired we set up a table that exhibits the pairings:

It is to the second table that McNemar's test can be applied. Notice that the sum of the numbers in the second table is 85—the number of pairs of siblings—whereas the sum of the numbers in the first table is twice as big, 170—the number of individuals. The second table gives more information than the first. The numbers in the first table can be found by using the numbers in the second table, but not vice versa. The numbers in the first table give only the marginal totals of the numbers in the second table.

Related tests

The binomial sign test gives an exact test for the McNemar's test.
The Cochran's Q test is an extension of the McNemar's test for more than two "treatments".
The Liddell's exact test is an exact alternative to McNemar's test.
The Stuart–Maxwell test is different generalization of the McNemar test, used for testing marginal homogeneity in a square table with more than two rows/columns.
The Bhapkar's test is a more powerful alternative to the Stuart–Maxwell test, but it tends to be liberal. Competitive alternatives to the extant methods are available.
The McNemar's test is a special case of the Cochran–Mantel–Haenszel test; it is equivalent to a CMH test with one stratum for the each of the N pairs and, in each stratum, a 2x2 table showing the paired binary responses.

Popular movies

The Hunger Games (film) - 2012 American dystopian action thriller science fiction-adventure film directed by Gary Ross and based on Suzanne Collins’s 2008 novel of the same name. It is the first insta...
untitled Captain Marvel sequel - part of Marvel Cinematic Universe....
Killers of the Flower Moon (film project) - Killers of the Flower Moon - film project in United States of America. It was presented as drama, detective fiction, thriller. The film project starred Leonardo Dicaprio, Robert De Niro. Director of...
Five Nights at Freddy's (film) - Five Nights at Freddy's - film published in 2017 in United States of America. Scenarist of the film - Scott Cawthon....

Popular books

Book of Revelation - The Book of Revelation is the final book of the New Testament, and consequently is also the final book of the Christian Bible. Its title is derived from the first word of the Koine Greek text: apok...
Book of Genesis - account of the creation of the world, the early history of humanity, Israel's ancestors and the origins...
Gospel of Matthew - The Gospel According to Matthew is the first book of the New Testament and one of the three synoptic gospels. It tells how Israel's Messiah, rejected and executed in Israel, pronounces judgement on ...
Michelin Guide - Michelin Guides are a series of guide books published by the French tyre company Michelin for more than a century. The term normally refers to the annually published Michelin Red Guide , the oldest...
Psalms - The Book of Psalms , commonly referred to simply as Psalms , the Psalter or "the Psalms", is the first book of the Ketuvim , the third section of the Hebrew Bible, and thus a book of th...
Ecclesiastes - Ecclesiastes is one of 24 books of the Tanakh , where it is classified as one of the Ketuvim . Originally written c. 450–200 BCE, it is also among the canonical Wisdom literature of the Old Tes...
The 48 Laws of Power - non-fiction book by American author Robert Greene. The book...

Popular television series

The Crown (TV series) - historical drama web television series about the reign of Queen Elizabeth II, created and principally written by Peter Morgan, and produced by Left Bank Pictures and Sony Pictures Tel...
Friends - American sitcom television series, created by David Crane and Marta Kauffman, which aired on NBC from September 22, 1994, to May 6, 2004, lasting ten seasons. With an ensemble cast sta...
Young Sheldon - spin-off prequel to The Big Bang Theory and begins with the character Sheldon...
Modern Family - American television mockumentary family sitcom created by Christopher Lloyd and Steven Levitan for the American Broadcasting Company. It ran for eleven seasons, from September 23...
Loki (TV series) - upcoming American web television miniseries created for Disney+ by Michael Waldron, based on the Marvel Comics character of the same name. It is set in the Marvel Cinematic Universe, shar...
Game of Thrones - American fantasy drama television series created by David Benioff and D. B. Weiss for HBO. It...
Shameless (American TV series) - American comedy-drama television series developed by John Wells which debuted on Showtime on January 9, 2011. It...