Published September 6, 2026

Research and Statistics for the NCE: The One Pass You Need (With a Sorting Drill)

By David Zimmerman · 9 min read · Clinical content

Research and program evaluation is the area most counseling students dread and the one that rewards a single focused pass: the NCE tests definitions, not calculation. Scales of measurement, the normal curve, reliability versus validity and their subtypes, threats to internal validity, designs, errors, and evaluation types cover almost every research item[1]. This post is that pass, with a sorting drill for the distinctions students most often mix up and an eight-item quiz. The classic sources are Campbell and Stanley on design[2] and the standard review texts[3].

Measurement

ConceptWhat to know
Scales (NOIR)Nominal (categories) → ordinal (rank) → interval (equal units, no true zero; IQ, temperature) → ratio (true zero; counts, time)
Central tendencyMean (sensitive to outliers), median (middle; use with skew), mode (most frequent). Positive skew: tail right, mean > median
Normal curve68% within ±1 SD, 95% within ±2, 99.7% within ±3
Standard scoresz (M 0, SD 1) · T (M 50, SD 10) · IQ/deviation (M 100, SD 15) · stanines (1–9, M 5, SD 2) · percentile rank (not equal-interval)
Standard error of measurementBand around an observed score reflecting unreliability; larger when reliability is lower
Norm- vs. criterion-referencedCompared with a group vs. compared with a standard (mastery)

Reliability vs. validity

Reliability is consistency; validity is whether the test measures what it claims. A test can be reliable and not valid; it cannot be valid without being reliable. The subtypes are the exam’s favorite matching items; sort them.

Sort: reliability, validity, or threat to internal validity?

Tap an item, then tap the bucket it belongs to. Tap a placed item to send it back.

Items (15 left)

Reliability

    Validity

      Threat to internal validity

        TermDefinition
        Reliability: test-retest · alternate forms · internal consistency (Cronbach's alpha, split-half, KR-20) · inter-raterConsistency across time, versions, items, and scorers
        Validity: content · criterion (concurrent, predictive) · construct (convergent, discriminant) · faceCoverage of the domain; relationship to a criterion now or later; measures the construct; looks right
        Internal validityThe intervention, not something else, caused the change
        External validityGeneralizes beyond the sample and setting
        Threats (Campbell & Stanley)History, maturation, testing, instrumentation, regression, selection, mortality, and interactions

        Designs and inference

        • True experimental: random assignment + control group. Quasi-experimental: no random assignment (intact groups). Single-subject: AB, ABAB (reversal), multiple baseline. Correlational / ex post facto: no manipulation.
        • Qualitative: phenomenology (lived experience), grounded theory (build theory from data), ethnography (culture), case study, narrative; trustworthiness via triangulation, member checking, thick description.
        • Hypothesis testing: null vs. alternative; alpha .05; p value; Type I (false positive) vs. Type II (false negative); power = 1 − β; effect size (Cohen’s d: .2 small, .5 medium, .8 large).
        • Statistics by question: t-test (two means), ANOVA (three or more means; post hoc tests), chi-square (frequencies), correlation (relationship), regression (prediction), meta-analysis (pooled effect sizes).
        • Sampling: random, stratified, cluster, convenience; larger samples reduce sampling error.

        Program evaluation and research ethics

        • Needs assessment before, formative during, summative / outcome after; process vs. outcome; cost-benefit.
        • Evidence-based practice integrates best research, clinical expertise, and client values; outcome measures (OQ-45, PHQ-9) track it in practice.
        • Ethics: IRB review, informed consent, confidentiality of data, no deception without justification and debriefing, authorship credit (ACA Section G)[4].

        Read the question for the verb

        “Consistent” means reliability. “Measures what it claims” means validity. “Predicts” means predictive validity. “Caused” needs an experiment. Most research items are vocabulary items in disguise.

        18 research topics, each with a question

        Research & Program Evaluation is one of nine CACREP areas in CaseSavvy's NCE library; the free sample topic shows the format. Full library from $19/month.

        See NCE prep

        Quick check

        Interactive

        Research and statistics on the NCE

        1 of 8

        A program evaluator randomly assigns clients to a new anxiety group or a waitlist and compares GAD-7 change. This is:

        Related

        Sources

        1. NBCC — National Counselor Examination (NCE) overview, candidate handbook, and content outline
        2. Campbell, D. T., & Stanley, J. C. (1963). Experimental and Quasi-Experimental Designs for Research. Houghton Mifflin.
        3. Rosenthal, H. (2017). Encyclopedia of Counseling (4th ed.). Routledge.
        4. American Counseling Association — ACA Code of Ethics (2014) and ethics resources

        Frequently asked questions

        What is the difference between reliability and validity?

        Reliability is consistency (across time, forms, items, and raters); validity is whether a test measures what it claims (content, criterion, construct). A test can be reliable without being valid, but not valid without being reliable.

        What is a Type I error?

        Rejecting a true null hypothesis: a false positive. Alpha (usually .05) is the accepted risk of it. A Type II error is failing to reject a false null, a false negative; power is the probability of avoiding it.

        Ready to practice?

        Start drilling NCMHCE-style questions for free — no credit card required.

        Start Free Practice →