9.3 Video Surveillance Design, Storage & Analytics
Key Takeaways
- Pick an operational goal first—detection, observation, recognition, or identification—then choose lens, resolution, lighting, and placement to support that goal.
- Johnson's cycles-on-target idea and commercial DORI language in IEC 62676-4 are planning heuristics, not a formula the PSP grades to three decimal places.
- Storage scales with camera count, bitrate, duty cycle, and retention; a worked 16-camera, 4 Mbps, 30-day example is about 21 TB raw.
- Analytics such as loitering and line crossing are filters for operators; they are not substitutes for lighting, aim, or a human review.
- Evidentiary export needs a native or documented file, hashes, player notes, and an audit trail—not a phone video of a monitor.
Operational Goals Before Pixel Counts
A video surveillance system (VSS) is not "cover the site with 4K." Domain 2 design starts with what a reviewer must be able to do with the image at a named location, at a named time of day, under named lighting.
Practitioners borrow language from two related traditions. Johnson's criteria (military imaging) described how many resolution cycles on a target you need to detect, recognize, or identify something in a cluttered scene. Commercial camera work often uses DORI language—detect, observe, recognize, identify—reflected in IEC 62676-4 planning practice. Neither tradition is a secret exam equation you should invent coefficients for. Both are a way to force the question: are we trying to notice that a person entered a yard, or to identify a known employee at a card reader?
Use the four goals as operational, not poetic:
- Detection — you can tell that something of interest is present or moving. Perimeter yards, roofs, and long fences often only need detection plus a patrol cue.
- Observation (sometimes called classification) — you can tell a person from a vehicle, or a person from an animal, and follow the action.
- Recognition — you can tell the type of object or a known class of person (uniformed officer versus visitor, box truck versus sedan) with useful confidence.
- Identification — you can distinguish a specific person or read a unique attribute the investigation actually needs (a face at a portal, a license plate where law and lighting allow).
If the cash-office door needs identification, a 120-degree lobby camera in the ceiling is the wrong tool even if it is 8 megapixels. Pixels that fall on the floor and the far wall do not land on the face. Field of view (FOV), mounting height, and the distance to target set pixels on target more than the marketing megapixel number.
| Goal | What a reviewer can do | Typical placement implication |
|---|---|---|
| Detection | Notice presence or motion | Wider FOV, longer range, analytics as a cue |
| Observation | Follow activity; person vs vehicle | Medium FOV; enough light to see limbs and paths |
| Recognition | Know the class or a familiar type | Tighter scene; stable aim; less compression on the subject |
| Identification | Specify an individual or unique detail | Narrow FOV at the portal; lighting on faces; higher bitrate on that stream |
Resolution, Compression, Frame Rate, Lighting, IR, FOV
Resolution is the sensor and encoded pixel grid. It is wasted if the lens is wide, the subject is distant, or compression destroys edges. Specify the scene, not only the camera.
Compression (H.264, H.265/HEVC, and vendor codecs) reduces bitrate. Group-of-pictures structures and aggressive quantization smear faces and license plates, especially at night. For evidentiary cameras, cap compression and consider a second higher-quality stream. "Smart codec" motion-based savings help storage; they also surprise you when a still loiterer barely updates.
Frame rate (FPS) must match the activity. Two frames per second may detect a parking lot at 03:00. A cashier dispute or a fast turnstile needs more frames so hands and objects are not a blur. High FPS without bitrate budget just forces heavier compression.
Lighting is still the cheapest resolution upgrade. Cameras need scene illumination, not only a bright dome LED that blinds the sensor. Infrared (IR) extends monochrome performance at night; it does not paint color, it can produce hot spots and moth-to-flame insects, and it fails if the IR reflector is the camera's own housing. Specify IR throw versus the detection distance. Exterior identification of faces at night often needs white light or a well-placed illuminator, plus glare control, not only IR.
Field of view is lens plus mount. A varifocal lens that was never aimed after the lift left will not match the drawing. Recheck FOV at commissioning with a person at the point of interest, not with a technician standing under the camera.
Privacy masking blocks windows, neighboring yards, and restroom doors in the live and recorded image. Masks are an ethics and legal control. They must hold through pan-tilt-zoom (PTZ) moves if the camera can look into a private space. A mask that exists only on the public view while the recorded stream is unmasked is a policy failure.
NVR, VMS, and a Storage Calculation You Can Show Your Work For
An NVR (network video recorder) is typically an appliance that records camera streams and provides playback. A VMS (video management system) is the software layer: users, permissions, maps, failover, analytics plug-ins, and multi-site views. Many products blur the line. Design issues are the same: who can watch live, who can export, where the bits live, and how long they live.
Bitrate is the storage lever. A simple continuous-recording estimate is:
Storage ≈ cameras × bitrate × seconds per day × retention days / 8
Worked example: 16 cameras at 4 megabits per second each, continuous, 30 days.
- Aggregate bitrate = 16 × 4 Mbps = 64 Mbps
- Bytes per second = 64 / 8 = 8 MB/s
- Per day = 8 MB/s × 86,400 s ≈ 691,200 MB/day ≈ 0.69 TB/day
- Thirty days ≈ 20.7 TB raw
Add RAID overhead, file-system spare, motion-versus-continuous differences, audio tracks, and a growth margin. Motion-only recording can cut that number dramatically in a quiet warehouse and barely cut it in a 24-hour lobby. If the exam stem gives cameras, bitrate, and days, multiply; do not invent a proprietary vendor constant.
Retention is a legal and investigative choice. Seven days may miss a slow-burn theft. Ninety days may be required by a contract and will triple the 30-day number in the example above. Write retention per camera class (perimeter detection vs cashier identification) instead of one site-wide guess.
| Design lever | What it changes | What it does not fix |
|---|---|---|
| Megapixels | Potential detail | Bad aim, glare, distance |
| Bitrate / codec | Storage and edge quality | A dark scene |
| FPS | Motion continuity | Identification at 200 feet with a wide lens |
| Retention days | Investigation window | A camera that never saw the door |
| IR illuminator | Night monochrome reach | Color identification, washed-out near-field |
Analytics as Tools, Not Magic
Loitering, line crossing, intrusion boxes, object-left, and similar rules can cue an operator or a patrol. They depend on stable scenes, correct perspective, and honest lighting. Wind-blown trees, headlights, insects on IR, and a PTZ that was left pointing at the sky will all "detect." Treat analytics as a filter that still needs a human or a verified secondary system (EACS grant, intercom, patrol).
Do not specify analytics to paper over the wrong operational goal. Line crossing on a 4-pixel-tall person is theater. Face analytics cannot identify someone the lens never resolved. Record the rule, the exclusion masks, and the expected nuisance rate the same way you would for an alarm sensor.
Chain of Evidence Export
When video may become evidence, export is a control. Prefer the native file (or a documented, lossless-enough container the VMS supports) plus a player or codec note so a third party can open it years later. Compute hashes (SHA-256 or the platform's documented function) of the exported object. Log who exported what, when, and why. Watermarks and on-screen timestamps help a jury; they do not replace a hash if someone can re-encode a clip.
Avoid the phone-pointed-at-a-monitor pattern. Avoid silent trimming with no audit. If you clip to the incident, export a defined buffer before and after and record that you clipped. Multi-camera incidents need a shared timeline. PTZ tours should show where the camera was looking; a zoomed identification shot with no context shot is incomplete.
Cyber hygiene sits beside evidence: unique passwords, network segmentation, disabled default accounts, and firmware support. A VMS on the open internet with the default admin password is not a recording plan.
Name families without copying tables: IEC 62676 for video surveillance systems, including DORI planning in part 4; EN national adoptions of that family; UL 3044 (CCTV equipment) appears on some North American listings; UL 827 does not make your NVR a central station. Use the names to read a specification. Use operational goals to design the camera.
A design narrative says cameras on the cash-office door must support identification of a known employee at the reader. In practitioner terms, that operational goal is closest to which imaging intent?
Sixteen cameras record continuously at 4 Mbps each. Ignoring RAID and file-system overhead, what is the approximate storage for 30 days?
When exporting video that may become evidence, which practice best protects chain of custody?