5.1 Dictionary Hierarchy & Architecture

Key Takeaways

  • Most CAT programs let the reporter rank active dictionaries, and job or case dictionaries are usually placed above the personal (main) dictionary so case-specific entries win.
  • A higher-ranked entry masks, but does not change, an identical outline in a lower-ranked dictionary, so a case name can override a common word without corrupting the personal dictionary.
  • Case dictionaries hold terms shared across a whole matter, while job dictionaries hold terms for one witness or session.
  • Merging a case dictionary into the personal dictionary makes temporary proper nouns permanent and can cause mistranslations in unrelated jobs.
  • RTF/CRE lets reporters move dictionaries between CAT systems, although some native features and formatting may not transfer cleanly.
Last updated: September 2026

5.1 Dictionary Hierarchy & Architecture

In modern stenographic reporting, Computer-Aided Transcription (CAT) software functions as an ultra-low-latency real-time compiler. Every physical chord struck on a steno writer generates an electronic data packet—typically two to six bytes encoding specific key depressions—that is transmitted instantaneously to the host computer. Within milliseconds, the CAT translation engine must ingest this raw chord, evaluate it against multiple active dictionary databases, parse multi-stroke combinations, resolve contextual dependencies, apply spacing and capitalization rules, and render polished English text across both the reporter's monitor and external client browsing screens.

To accomplish this feat without lag or ambiguity, CAT software enforces a strict dictionary lookup hierarchy. Understanding how this hierarchical cascade operates—and how to manipulate it to protect your personal master dictionary while capturing case-specific jargon—is essential knowledge for the NCRA Written Knowledge Test (WKT) and the foundation of professional realtime competency.


A Typical CAT Dictionary Lookup Order

When the translation engine receives a steno stroke or stroke sequence, it does not scan dictionary databases at random. Instead, it checks the active dictionaries in the order the reporter has set (exact behavior varies by CAT program, so learn your own system's settings). The working rule is:

Higher Priority DictionaryLower Priority Dictionary\text{Higher Priority Dictionary} \succ \text{Lower Priority Dictionary}

If an identical steno outline exists in multiple active dictionaries, the translation definition residing in the dictionary with the highest priority rank is rendered on screen. Lower-ranked definitions for that exact same outline are ignored for that instance, remaining entirely suppressed until the higher-ranking dictionary is unloaded or unlinked.

┌─────────────────────────────────────────────────────────────────────────┐
│                     CAT TRANSLATION LOOKUP CASCADE                      │
├─────────────────────────────────────────────────────────────────────────┤
│  1. JOB / CASE DICTIONARIES        (Highest Priority - Temporary/Local) │
│     ▲ Case captions, expert names, local street names, trade acronyms   │
│     │                                                                   │
│  2. PERSONAL / MAIN DICTIONARY     (Core Lifelong Operating Layer)      │
│     ▲ 80,000–250,000+ entries: vocabulary, briefs, punctuation, speakers│
│     │                                                                   │
│  3. SPECIALTY / DOMAIN GLOSSARIES  (Secondary Layer - Subject Specific) │
│     ▲ Dorland's Medical, Patent Jargon, Maritime, Construction Defects  │
│     │                                                                   │
│  4. SYSTEM / SPELLING DICTIONARIES (Base Vendor Tables - English Corpus)│
│     ▲ Standard vendor wordlists, phonetic fallbacks, spelling tables    │
│     │                                                                   │
│  5. AFFIX & SUFFIX PARSING ENGINES (Morphological Rule Folders)         │
│     ▼ Automated compound assembly: un-, re-, -ing, -ed, -s, -ment       │
└─────────────────────────────────────────────────────────────────────────┘

1. Job and Case Dictionaries (Layer 1 - Highest Precedence)

The Job Dictionary (specific to a single witness or session) and Case Dictionary (shared across an entire multi-day or multi-month litigation) occupy the apex of the lookup pyramid. When counsel or a witness speaks an unusual proper noun, foreign surname, or proprietary corporate acronym, that term is defined directly into the active case dictionary. Because this layer is evaluated first, it immediately supersedes any conflicting outlines in the reporter's general vocabulary.

2. Personal / Main Dictionary (Layer 2 - Core Operating Layer)

The reporter's Personal Dictionary (often called the Master Dictionary) is their primary professional asset. Built and refined across years or decades of daily writing, this file typically contains between 80,000 and 250,000+ unique entries. It encapsulates the reporter's personal steno theory, customized briefing systems, number formatting shortcuts, punctuation definitions, and speaker tokens (e.g., THE COURT:, THE WITNESS:). The personal dictionary translates the vast majority of standard English speech during any proceeding.

3. Specialty / Domain Dictionaries (Layer 3 - Secondary Glossaries)

Reporters frequently encounter proceedings governed by dense, highly technical vocabularies. Rather than permanently bloating their main dictionary with thousands of obscure terms, reporters link secondary specialty dictionaries—such as medical terminology (e.g., Stedman's or Dorland's), pharmaceutical trade registers, patent engineering glossaries, or environmental litigation tables. These dictionaries are queried only when an outline is absent from both the case and personal dictionaries.

4. System / Spelling Dictionaries (Layer 4 - Base Vendor Corpus)

Supplied by the CAT software vendor (e.g., Stenograph, Advantage Software, ProCAT, Gigatron), the system dictionary contains hundreds of thousands of standard English words and root forms. It acts as an underlying linguistic net, ensuring that multi-syllabic words stroked according to textbook phonetic theory can translate even if the reporter has never manually entered them into their personal dictionary.

5. Affix & Suffix Parsing Engines (Layer 5 - Morphological Fallbacks)

If a complete word outline is not found in Layers 1 through 4, the CAT engine attempts morphological assembly. It looks for known root words in active dictionaries and tests whether trailing or leading strokes match programmed prefix (e.g., un-, re-, anti-) or suffix (e.g., -ing, -ed, -tion, -ment) rules.


Priority Precedence in Action: Preventing Translation Disasters

The practical value of this strict hierarchy is best illustrated through real-world examples where common nouns collide with case-specific proper names:

Steno ChordPersonal Dictionary Translation (Layer 2)Case Dictionary Translation (Layer 1)Active Proceeding OutputRationale & Protection
PHEUL/ERmiller (common noun)Mr. Miller (plaintiff counsel)Mr. MillerThe case dictionary overrides the common noun for the duration of the trial without altering the reporter's personal definition.
PWAEU/K-Rbaker (profession/noun)Dr. Baker (testifying surgeon)Dr. BakerEliminates manual capitalization and prefix insertion during expert testimony while preserving lowercase "baker" for general use.
K-R-Bcurb (concrete street border)KIRB (proprietary software system)KIRBEnables lightning-fast realtime output of a technical acronym without creating permanent translation conflicts in future traffic accident cases.
SKWR-PBgene (biological DNA unit)Jean (co-defendant)JeanSolves a direct homophonic conflict for a specific criminal trial where the co-defendant's first name matches a standard biological noun.

[!IMPORTANT] The Override Principle: A higher-priority dictionary never deletes, overwrites, or modifies entries in a lower-priority dictionary. It merely masks or supersedes them during active translation. When the case dictionary is detached at the conclusion of the trial, the reporter's personal dictionary immediately re-emerges with all its original definitions intact and uncorrupted.


Creating and Scoping Job-Specific Dictionaries

Professional court reporters prepare for complex litigation by systematically harvesting specialized terminology prior to the start of proceedings. This preparation process, known as scoping and dictionary pre-population, follows a structured workflow:

[Notice of Deposition / Court Docket / Pleadings / Exhibits]
                           │
                           ▼
        [Extract Proper Nouns, Chemical Names & Acronyms]
                           │
                           ▼
     [Determine Theory-Compliant, Non-Conflicting Steno Outlines]
                           │
                           ▼
        [Compile Directly into Scoped Case Dictionary]
                           │
                           ▼
     [Attach to Active Realtime CAT Session Above Main Dictionary]

Sources for Pre-Trial Vocabulary Harvesting

  1. Case Pleadings & Dockets: Notice of deposition, complaint, answer, witness lists, exhibit indices, and joint pre-trial statements contain party names, attorneys of record, addresses, and statutory citations.
  2. Expert Witness Curricula Vitae (CVs): Medical degrees, board certifications, specialized procedural terminology, research publications, and professional affiliations.
  3. Patent Claims & SEC Filings: Detailed technical diagrams, patent claim numbers, pharmaceutical molecular names, and corporate subsidiaries.
  4. Local Geographic Datasets: Regional street maps, subdivision names, municipal landmarks, and highway intersections relevant to accident reconstruction.

Scoping Boundaries: Case vs. Job Dictionaries

In large-scale Multi-District Litigation (MDL) or high-stakes commercial trials involving dozens of depositions across multiple reporting teams, dictionary architecture must be carefully partitioned:

  • Case Dictionaries: Broad scope. Contains terms common to all witnesses in the entire dispute (e.g., corporate entity names, primary patent numbers, lead counsel names, aircraft component part numbers). Shared among all reporters working on the litigation.
  • Job Dictionaries: Narrow scope. Created specifically for an individual witness or single session (e.g., Dr. Smith's clinic name, his specific surgical patent, his assistant's surname). Loaded only when that specific witness is on the stand.

[!WARNING] The "Merge All" Disaster: A catastrophic error committed by novice reporters is using the automated "Merge Case Dictionary to Main Dictionary" function at the conclusion of a proceeding. Merging a 500-word case dictionary into your lifelong master dictionary permanently contaminates your primary vocabulary with case-specific proper nouns, causing bizarre mistranslations in subsequent, unrelated litigation for months or years to come.


Dictionary File Architecture, RTF/CRE & Memory Management

To sustain real-time performance at 225+ words per minute, CAT software relies on highly optimized database structures and operating system memory management.

Native Dictionary Files vs. RTF/CRE

Each CAT program stores dictionaries in its own native format; Stenograph's documentation, for example, identifies Case CATalyst dictionaries as .sgdct files. Native formats are built for fast lookup but cannot be opened by competing programs. To move a dictionary between systems, reporters export it to RTF/CRE, where each entry pairs a steno outline in a \cxs control word with its translation:

{\rtf1\ansi{\*\cxrev100}\cxdict{\*\cxsystem Case CATalyst}
{\*\cxs KAT}cat\par
{\*\cxs TKOG}dog\par
}

RTF/CRE is the common exchange route, but it is not always lossless: vendor-specific commands, conflict markers, and some formatting may be dropped or converted, so test an imported dictionary before relying on it in realtime.

Multi-Dictionary Layering and RAM Optimization

CAT software lets reporters link several dictionaries to an active realtime session. When a session opens, the CAT software loads the index tables of all active dictionaries directly into the computer's Random Access Memory (RAM) cache. Keeping lookups in memory avoids waiting on the disk, even a fast SSD.

[Physical Storage: NVMe SSD] ──▶ (Startup Load) ──▶ [RAM Cache: 64-Bit Memory]
                                                            │
[Steno Writer: RS-232/USB]   ──▶ [Translation Engine] ◀─────┘
                                         │ (Sub-Millisecond Lookup)
                                         ▼
                            [Instant Realtime Screen Output]

Realtime Performance Risks and Latency Traps

Improper dictionary management can degrade translation speed, resulting in "realtime stutter" or lagging text on client screens:

  • Corrupt Index Headers: Abrupt system crashes or forced reboots during an open session can corrupt dictionary index pointers, causing the CAT engine to hang while searching damaged nodes.
  • Excessive Unindexed Files: Activating dozens of uncurated, highly fragmented secondary dictionaries forces the translation loop to search deeper chains, increasing CPU thread utilization.
  • Routine Maintenance: Many CAT programs include dictionary maintenance tools that find duplicates, conflicts, and damaged entries; run them periodically and before large jobs.
Test Your Knowledge

In a typical CAT setup, which dictionary order lets case-specific entries override the reporter's general vocabulary?

A
B
C
D
Test Your Knowledge

Why should specialized proper nouns, technical acronyms, and expert witness names be placed into a scoped Case or Job dictionary rather than the reporter's permanent Personal dictionary?

A
B
C
D
Test Your Knowledge

What is the primary role of the RTF/CRE standard in stenographic dictionary and transcript architecture?

A
B
C
D