8.1 Evaluate the Argument: The Variance Test and Decisive Information

Key Takeaways

  • Evaluate the Argument questions ask test-takers to identify the single piece of missing information or empirical inquiry that is most critical to assessing the validity of a conclusion or the viability of a proposed plan.
  • The Variance Test is the definitive operational tool for Evaluate questions: test the two polar extreme answers to each inquiry (such as 'Yes' vs. 'No' or 100% vs. 0%); the correct answer choice will decisively strengthen the argument at one extreme and decisively undermine it at the other.
  • Incorrect answer choices fail the Variance Test because their extreme answers produce either no logical impact on the conclusion, produce identical impacts in both directions, or address secondary background context rather than the author's core unstated assumption.
  • Evaluate stimuli predominantly center on business plans, policy interventions, cost-benefit predictions, and causal hypotheses; locating the central gap between the premises and the conclusion reveals the missing variable.
  • Common traps include one-sided relevance (an inquiry that matters only under a pre-existing bias), out-of-scope cross-industry comparisons, and pseudo-evaluative options that merely re-verify facts already confirmed in the stimulus.
Last updated: September 2026

8.1 Evaluate the Argument: The Variance Test and Decisive Information

Quick Summary: "Evaluate the Argument" questions belong to the broader Assumption/Strengthen/Weaken logical family. Instead of asking you to select a premise that directly helps or harms the argument, an Evaluate stem asks what inquiry would be most useful to determine whether the argument's conclusion holds. The ultimate validation technique is the Variance Test: supplying the two opposite extreme answers to an option must cause the argument's validity to bifurcate—one extreme validates the conclusion, while the opposite extreme shatters it.

On the GMAT Focus Edition, Critical Reasoning questions test your capacity to dissect reasoning under executive decision-making conditions. Business leaders rarely receive perfectly complete data sets; instead, they must evaluate which additional data streams are essential before committing capital to a project or strategy. Evaluate the Argument questions replicate this precise managerial challenge.


The Architecture of Evaluate Questions

Evaluate questions present an argument with an inferential leap: premises are provided, and an author draws a definite conclusion or recommends a specific course of action. However, the author has relied upon an unstated assumption. Without knowing the factual reality of that assumption, the argument remains vulnerable.

Typical Question Stems

  • "Which of the following would it be most useful to determine in evaluating the argument?"
  • "The answer to which of the following questions would be most important in assessing the likelihood that the plan will succeed?"
  • "Which of the following investigations would be most critical in evaluating the validity of the analyst's conclusion?"
  • "Which of the following pieces of information would be most decisive in determining whether the city council should adopt the proposed policy?"

Logical Relationship to Strengthen and Weaken Questions

An Evaluate question is effectively a two-way Strengthen/Weaken question in neutral disguise:

  • In a Strengthen question, GMAC supplies the positive extreme of the assumption as an established fact.
  • In a Weaken question, GMAC supplies the negative extreme of the assumption as an established fact.
  • In an Evaluate question, GMAC poses the underlying variable as an open inquiry (e.g., "Whether $X$ is true or false"). If the answer is "Yes," the conclusion is strengthened; if the answer is "No," the conclusion is weakened.
                    [ Unstated Assumption / Core Inferential Gap ]
                                      │
          ┌───────────────────────────┴───────────────────────────┐
          ▼                                                       ▼
[ Strengthen Question ]                                 [ Weaken Question ]
Asserts the Assumption is True                          Asserts the Assumption is False
          │                                                       │
          └───────────────────────────┬───────────────────────────┘
                                      ▼
                         [ Evaluate Question ]
                   "Whether the Assumption is True"
              Polar Extreme A (+) ──▶ Strengthens Conclusion
              Polar Extreme B (-) ──▶ Weakens Conclusion

The Operational Tool: The Variance Test

The Variance Test is the gold standard technique for verifying the correct answer choice on high-difficulty Evaluate questions. When down to two tempting options, applying the Variance Test eliminates ambiguity.

Step-by-Step Variance Test Methodology

  1. Isolate the Core Conclusion or Plan: Pinpoint the exact claim made by the author. Clarify the precise scope, metric of success (e.g., net profitability, reduction in traffic accidents, patient survival rate), and temporal frame.
  2. Identify the Inferential Gap: Formulate the central unstated assumption bridging the premises to that conclusion.
  3. Identify the Polarity of the Answer Choice: Most choices begin with phrases like "Whether...", "The extent to which...", "The proportion of...", or "If...".
  4. Inject Polar Extreme Values: Test the two furthest extremes of the inquiry:
    • For a Whether inquiry: test Definite "Yes" vs. Definite "No".
    • For a Proportion / Extent inquiry: test 100% (All) vs. 0% (None).
    • For a Comparative Cost / Price inquiry: test Substantially Exceeds vs. Negligible / Zero.
  5. Evaluate the Bifurcation Impact:
    • The Correct Choice: One extreme must significantly bolster the argument, while the opposite extreme must substantially undermine or destroy it.
    • A Distractor (Incorrect Choice): Both extremes leave the argument in virtually the same state (yielding a "So what?"), or the inquiry affects an ancillary fact that does not touch the inferential leap between the premises and the conclusion.

Variance Test Demonstration Matrix

Candidate InquiryExtreme 1 ("Yes" / High)Impact on ArgumentExtreme 2 ("No" / Low)Impact on ArgumentVerdict
"Whether operating costs will exceed projected energy savings"Yes, operating costs will dwarf savingsDestroys profitability claimNo, operating costs are negligibleConfirms profitability claimCORRECT (Decisive Leverage)
"Whether the equipment manufacturer produces other machinery"Yes, they make dozens of product linesNeutral ("So what?")No, they make only this machineNeutral ("So what?")INCORRECT (Out of Scope)
"Whether competitor firms in other countries use similar systems"Yes, international firms use itWeak analogy; does not prove local outcomeNo, international firms do notStill depends on local cost structureINCORRECT (Collateral Detail)

Applied Frameworks: Business Plans vs. Causal Arguments

Evaluate stimuli typically fall into one of two structural archetypes:

1. The Business Plan / Policy Strategy Archetype

  • Structure: The author proposes a plan $P$ to achieve goal $G$ (usually increasing profit, reducing cost, improving productivity, or reducing risk).
  • Underlying Assumptions:
    • The implementation of $P$ will not trigger unexpected offsetting expenses or counter-productive side effects.
    • Consumers or target agents will respond as forecasted rather than circumventing the initiative.
    • External market conditions will remain sufficiently stable during the execution horizon.
  • Variance Test Target: Inquire whether the hidden costs or behavioral counter-responses exist in sufficient magnitude to overwhelm the planned benefit.

2. The Causal Hypothesis Archetype

  • Structure: An observed correlation between $X$ and $Y$ leads the author to conclude that $X$ caused $Y$.
  • Underlying Assumptions:
    • An unmeasured third factor $Z$ did not cause both $X$ and $Y$.
    • The causal direction is not reversed ($Y$ did not cause $X$).
    • The observed correlation is not an artifact of data collection bias or statistical noise.
  • Variance Test Target: Inquire whether factor $Z$ was present during the test period, or whether $Y$ preceded $X$.

Worked Analytical Walkthroughs

Worked Example 1: Corporate Operations & Fleet Electrification

Stimulus:
Metro Haulage operates a fleet of 500 diesel delivery trucks in a major metropolitan area. Seeking to lower operating expenses, the management plans to replace the entire diesel fleet with newly developed electric commercial vans over the next two years. Management argues that this transition will significantly improve Metro Haulage's net operating profitability because electric vans consume electricity that costs roughly 60 percent less per mile than diesel fuel, and electric drivetrains require far fewer scheduled mechanical maintenance procedures.

Question Stem:
Which of the following would it be most useful to determine in evaluating whether Metro Haulage's plan will achieve its financial objective?

Options:
(A) Whether other logistics companies in different geographic regions have successfully incorporated electric vehicles into their fleets.
(B) Whether the residual resale value of the existing diesel trucks will remain stable over the next five years.
(C) Whether the capital depreciation and battery-replacement costs associated with the new electric vans will exceed the projected fuel and routine maintenance savings over the operating life of the vehicles.
(D) Whether local municipality regulations will impose additional restrictions on diesel truck emissions in urban centers during the coming decade.
(E) Whether the drivers employed by Metro Haulage prefer the handling characteristics of electric commercial vans over diesel trucks.

Analytical Deconstruction

  1. Identify the Plan and Goal:
    • Plan: Replace 500 diesel delivery trucks with electric commercial vans.
    • Goal: Significantly improve net operating profitability.
    • Premises: Electricity costs 60% less per mile than diesel; electric drivetrains require fewer routine maintenance procedures.
  2. Isolate the Gap: Lower fuel and routine maintenance costs per mile do not guarantee higher net operating profitability if other vehicle costs (such as capital amortization, replacement batteries, or charging infrastructure) outweigh those operational savings.
  3. Apply the Variance Test:
    • Consider the inquiry in choice (C): "Whether capital depreciation and battery-replacement costs will exceed projected fuel and routine maintenance savings."
      • Extreme 1 (YES): Yes, battery-replacement and depreciation costs are enormous, totaling $1.20 per mile compared to fuel/maintenance savings of $0.40 per mile. Net operating expenses rise, and profitability declines. The plan fails (Massive Weaken).
      • Extreme 2 (NO): No, battery replacements are rare and covered under long-term warranties; capital costs are modest, taking up only $0.10 per mile while savings are $0.40 per mile. Net operating expenses fall sharply. The plan succeeds (Strong Strengthen).
    • Because the two polar outcomes produce opposite, decisive effects on the conclusion, this inquiry is logically indispensable.
  4. Evaluate Distractors:
    • Choice (A) addresses other companies in different geographic regions; differing geography, electricity rates, and routes render this comparison inconclusive.
    • Choice (B) discusses the resale value of diesel trucks over five years; while trade-in value provides a one-time cash inflow, the conclusion concerns the operating profitability of the electric fleet over time.
    • Choice (D) discusses future diesel emission regulations; future restrictions on diesel might penalize diesel, but they do not establish whether the electric fleet itself will be profitable.
    • Choice (E) concerns driver subjective preferences; driver preference does not determine financial viability unless linked to wage demands or retention, which is not indicated.

Worked Example 2: Public Health & Causal Inference

Stimulus:
Over the past five years, the incidence of severe dental cavities among school-aged children in the township of Pinecrest dropped by 35 percent. Five years ago, the township installed an advanced fluoridation system in its municipal water reservoir. Health officials concluded that the municipal water fluoridation was directly responsible for the dramatic reduction in childhood dental cavities.

Question Stem:
Which of the following investigations would be most important in evaluating the health officials' causal conclusion?

Options:
(A) Determining whether adult residents of Pinecrest experienced a similar 35 percent drop in cavity rates over the same five-year period.
(B) Determining whether neighboring townships with similar socioeconomic demographics that did not fluoridate their water also experienced a comparable decline in childhood cavity rates over the past five years.
(C) Determining whether the municipal fluoridation system required extensive repairs or filtration upgrades during its initial year of operation.
(D) Determining the exact chemical concentration of fluoride maintained in Pinecrest's municipal water supplies compared to federal guidelines.
(E) Determining whether pediatric dentists in Pinecrest recently adopted newer composite resin materials for filling existing tooth cavities.

Analytical Deconstruction

  1. Identify the Causal Conclusion: Municipal water fluoridation caused the 35% drop in childhood cavities in Pinecrest.
  2. Identify the Gap: The officials observe a correlation (fluoridation installed 5 years ago; cavity rates fell over 5 years). They assume no other simultaneous regional change (such as school dental programs, widespread sealants, changes in dietary sugar, or a broader regional trend) caused the decline.
  3. Apply the Variance Test to Choice (B):
    • Inquiry: Did neighboring non-fluoridated townships experience a comparable decline?
      • Extreme 1 (YES): Neighboring townships without fluoridation also saw cavity rates drop by 35%. This demonstrates that the decline occurred independently of water fluoridation (due to regional dietary shifts, improved fluoride toothpaste, or school programs). The causal claim is destroyed (Weaken).
      • Extreme 2 (NO): Neighboring townships without fluoridation saw cavity rates remain flat or increase. This isolated control group proves that Pinecrest's decline was unique, strongly isolating fluoridation as the causal driver. The causal claim is verified (Strengthen).
  4. Evaluate Distractors:
    • Choice (A) investigates adult residents; adult teeth differ from developing childhood teeth in susceptibility to fluoridation, so adult outcomes do not resolve the pediatric causal link.
    • Choice (C) asks about filtration repairs during year one; short-term maintenance does not explain a sustained five-year outcome.
    • Choice (D) asks about federal guidelines; knowing compliance does not isolate causation from external confounding factors.
    • Choice (E) focuses on composite resin filling materials; filling materials treat cavities after they exist, whereas the stimulus measures the incidence (onset) of cavities.

High-Frequency GMAT Traps on Evaluate Questions

Trap 1: The One-Sided Relevance Distractor

A distractor often provides information that could weaken the argument if the answer turns out to be "Yes," but if the answer is "No," the argument is unaffected (or vice versa). A true evaluation item must provide significant evaluative leverage in both directions.

Trap 2: The Pseudo-Evaluative Re-Statement

Some distractors ask whether a premise already provided in the stimulus is accurate (e.g., "Whether electricity actually costs 60% less than diesel"). On the GMAT, premises are taken as factual. You do not evaluate an argument by questioning the truth of the stated premises; you evaluate the validity of the inferential leap connecting those premises to the conclusion.

Trap 3: Irrelevant Relative Comparison to External Groups

Options frequently invite you to compare the subject company or municipality to an outside entity (e.g., "Whether competitors spend more on advertising" or "Whether another state has higher tax rates"). Unless the conclusion explicitly asserts a comparative superiority over those external entities, external benchmarks are out of scope.

Trap 4: Vague Quantifiers That Fail the Variance Test

Be wary of answer choices containing words like "some," "a few," or "occasionally." For instance, knowing whether "some drivers experience fatigue" has virtually no evaluative leverage because "some" can mean as few as two drivers, which is mathematically insufficient to alter an institutional fleet plan.

Loading diagram...
The Operational Variance Test Funnel for Evaluate Questions
Test Your Knowledge

A national retail pharmacy chain plans to install automated self-checkout kiosks across all 1,200 of its retail stores over the next 18 months. Executive leadership argues that this deployment will substantially increase the chain's operating profits by allowing individual store managers to reduce the number of paid cashier labor hours by 40 percent without causing customer wait times at checkout to increase. Which of the following would it be most useful to determine in evaluating whether the pharmacy chain's plan will achieve its projected financial objective?

A
B
C
D
Test Your Knowledge

Agricultural researchers observed that wheat farms in the Red River Valley that utilized a novel organic mycorrhizal soil treatment suffered 60 percent less crop damage from seasonal fungal root rot than did adjacent conventional farms that used synthetic fungicides. The researchers concluded that applying this mycorrhizal soil treatment is more effective at protecting wheat crops against fungal root rot than applying synthetic fungicides. Which of the following investigations would be most important in evaluating the agricultural researchers' conclusion?

A
B
C
D
Test Your Knowledge

The municipal transit authority of Riverdale recently replaced the flat-rate $2.50 transit fare with a dynamic distance-based pricing model under which short-distance trips cost $1.50 and long-distance suburban trips cost $4.00. The transit commissioner predicts that this structural fare adjustment will increase total annual transit revenue because the lower short-distance fare will attract thousands of urban commuters who currently walk or ride bicycles for short errands. Which of the following would it be most useful to establish in order to evaluate the transit commissioner's prediction?

A
B
C
D