16.2 Configuration Management: Comparing Ansible, Puppet, Chef, and SaltStack
Key Takeaways
Configuration management frameworks eliminate snowflake configurations, prevent configuration drift, and replace error-prone manual CLI changes with declarative, auditable, and version-controlled workflows.
Ansible operates using an agentless architecture over standard SSH or NETCONF/RESTCONF APIs, using YAML playbooks and Python execution modules on the control node.
Puppet and Chef historically employ agent-based, pull-oriented models with domain-specific languages (Puppet DSL and Ruby); when managing network switches that cannot run native agent daemons, they utilize proxy agents or external device drivers.
SaltStack features a high-speed, event-driven architecture utilizing a ZeroMQ message bus and Python Minion agents, while supporting agentless network automation via Salt-SSH and Salt-Proxy.
Idempotency guarantees that executing an automation playbook or catalog multiple times against a device produces the identical desired target state without creating unintended duplicate configurations or disruptive side effects.
Configuration Management: Comparing Ansible, Puppet, Chef, and SaltStack
Traditional enterprise network administration relies heavily on manual Command-Line Interface (CLI) configuration sessions over SSH. While practical for small-scale deployments, manual management becomes unsustainable in complex enterprise campuses, datacenters, and software-defined WANs. Manual administration introduces significant operational risks that threaten network availability and governance.
The Operational Imperative for Configuration Management
Configuration management tools modernize network operations by resolving four fundamental challenges:
- Snowflake Configurations: Over months and years, individual switches and routers receive bespoke, undocumented CLI adjustments during troubleshooting sessions or emergency maintenance. These devices become unique "snowflakes" that diverge from baseline architecture standards, making upgrades and troubleshooting unpredictable.
- Configuration Drift: The gradual divergence between the actual running configuration on physical devices and the intended gold-standard design documented in central repositories. Drift often occurs when out-of-band changes bypass standard change management controls.
- Human Error: Manual typing mistakes, transposed subnet octets, misplaced access-list entries, and forgotten security hardening commands are a leading cause of enterprise network outages.
- Lack of Auditability and Version Control: Direct CLI interactions leave minimal audit trails beyond fragmented syslog records. Version-controlled configuration management provides a complete Git commit history detailing who changed what, when, and why.
+-------------------------------------------------------------------------+
| Configuration Management Architectural Paradigms |
+-------------------------------------------------------------------------+
| |
| AGENTLESS ARCHITECTURE (Ansible) |
| +--------------+ SSH / NETCONF / RESTCONF +-------------+ |
| | Control Node | ==================================> | Switch/Rtr | |
| | (Python/YAML)| | (No Agent) | |
| +--------------+ +-------------+ |
| |
| AGENT-BASED ARCHITECTURE (Puppet / Chef / SaltStack Server Models) |
| +--------------+ Periodic Pull (Poll) +-------------+ |
| | Master Node | <---------------------------------- | Node/Agent | |
| | (Catalog/DSL)| ==================================> | (Daemon) | |
| +--------------+ State Convergence +-------------+ |
| |
| NETWORK PROXY MODEL (Managing Non-Linux Network Appliances) |
| +--------------+ Catalog +---------------+ SSH/API +--+ |
| | Master Node | =================> | Proxy Agent | =========> |Sw| |
| | | | (Device Proxy)| +--+ |
| +--------------+ +---------------+ |
+-------------------------------------------------------------------------+
Architectural Comparison: Agentless vs. Agent-Based
A primary architectural distinction among configuration management platforms is whether target endpoints require specialized background daemon software:
Agentless Architecture (Ansible)
Ansible operates without installing any custom agent software on managed endpoints. Instead, the Ansible control machine connects to target devices using standard transport protocols—primarily SSH for CLI or NETCONF, and HTTPS for RESTCONF and REST APIs. The control machine executes Python modules locally or transmits commands directly to the endpoint. This architecture is ideal for network infrastructure, where closed operating systems (such as Cisco IOS, IOS-XE, and NX-OS) prevent administrators from running arbitrary background daemons.
Agent-Based Architecture (Puppet, Chef, SaltStack)
Traditional server automation platforms rely on a resident software agent (such as puppet-agent, chef-client, or salt-minion) installed inside the operating system of the target endpoint. The agent periodically communicates with the master server, collects local facts, receives configuration instructions, and enforces state locally. Because network switches have proprietary microcode or restricted control planes, agent-based tools adapt using specialized proxy architectures:
- Puppet: Utilizes
puppet-device, an intermediary proxy daemon running on an external server (or within a switch's Linux Guest Shell / container) that translates declarative Puppet catalogs into device-specific CLI, NETCONF, or RESTCONF transactions. - Chef: Utilizes Chef Network and proxy drivers (such as the Train transport library) to bridge the central Chef Server to remote switch CLIs or APIs.
- SaltStack: Employs
salt-proxy, a persistent proxy minion running on an external node that maintains stateful SSH or NETCONF sessions with network appliances, or leveragessalt-sshfor agentless execution.
Communication Paradigms: Push vs. Pull Models
Configuration management platforms distribute configurations using either push or pull models:
Push Model (Ansible, SaltStack)
In a push architecture, the centralized control node initiates communication on demand. When an engineer or CI/CD pipeline triggers execution, the control server connects out to target devices and delivers the configuration immediately. This model provides:
- Immediate, deterministic execution timing.
- Coordinated multi-device changes, essential when updating routing protocols across redundant core switches (e.g., ensuring both HSRP/VRRP peers or BGP neighbors converge simultaneously).
- Simpler device management, as endpoints do not require outbound connectivity to a master server.
Pull Model (Puppet, Chef)
In a pull architecture, managed agents initiate communication on a periodic schedule (e.g., every 30 minutes). The local agent wakes up, contacts the central master server, requests its compiled configuration catalog, compares it against the local running state, and converges any detected discrepancies. This model excels at:
- Continuous, automated drift correction without human intervention.
- Massive horizontal scalability across thousands of distributed compute nodes.
- However, pull-based scheduling can cause timing mismatches during synchronized network topology migrations.
Data Serialization and Domain Languages
The choice of configuration language impacts maintenance overhead and required team skillset:
- Ansible: Uses YAML (YAML Ain't Markup Language) for playbooks, combined with Jinja2 templating for dynamic configuration generation. Playbooks call Python execution modules under the hood, requiring no programming expertise from network engineers.
- Puppet: Uses a declarative Puppet Domain-Specific Language (DSL) based on Ruby syntax, organizing resources into manifests (
.ppfiles) and modules. - Chef: Uses pure Ruby syntax for defining configuration rules organized into recipes and cookbooks, offering programmatic flexibility but requiring Ruby programming proficiency.
- SaltStack: Uses YAML and Jinja2 for state definition files (
.sls), structured execution modules, and a hierarchical data repository called Pillar for managing secrets and device-specific variables.
Master Comparison Matrix
| Dimension | Ansible | Puppet | Chef | SaltStack |
|---|---|---|---|---|
| Primary Architecture | Agentless (control node connects directly) | Master-Agent (Puppet Master / Agent) | Client-Server (Chef Server / Client) | Master-Minion / Event-driven Bus (ZeroMQ) |
| Network Device Strategy | Direct SSH, NETCONF, RESTCONF, REST APIs | puppet-device proxy agent or Guest Shell | Chef Network proxy / Train transport | salt-proxy minion or salt-ssh |
| Communication Model | Push (immediate control node dispatch) | Pull (default 30-min agent query); Push capable | Pull (default periodic agent query); Push capable | Push (ultra-fast ZeroMQ publish) or Event-driven |
| Configuration Syntax | YAML playbooks with Jinja2 templating | Puppet DSL (declarative Ruby-like) | Pure Ruby scripts (cookbooks/recipes) | YAML state files (.sls) with Jinja2 |
| Core Language | Python | Ruby / C++ | Ruby | Python |
| Learning Curve | Low (clean declarative YAML) | Moderate (specialized DSL) | High (requires Ruby programming proficiency) | Moderate (YAML with advanced Jinja2 logic) |
| State Management | Stateless; compares device state dynamically | Compiled catalog compared by agent | Chef Server resource collection / node object | High-speed cache and persistent Pillar data |
Understanding Idempotency in Network Automation
Idempotency is a foundational principle of modern infrastructure automation. An operation is defined as idempotent if executing it multiple times yields the exact same system state as executing it once, without producing unintended side effects, duplicate entries, or unnecessary state churn.
Non-Idempotent (Imperative) vs. Idempotent (Declarative) Execution
Consider an imperative shell script appending an NTP server to a device:
# Non-Idempotent Imperative Script:
ssh admin@10.1.1.1 "configure terminal ; ntp server 192.168.1.100 ; end ; write memory"
If executed three times, some operating systems might add three identical lines or re-initialize the NTP synchronization daemon unnecessarily, causing transient clock skew. In contrast, an idempotent Ansible task evaluates current state against desired state:
# Idempotent Declarative Task:
- name: Configure Enterprise NTP Server
cisco.ios.ios_ntp_global:
config:
servers:
- server: 192.168.1.100
prefer: true
state: merged
Idempotent Module Execution Workflow
- Query Current State: The module connects to the device and retrieves running configuration data via SSH CLI or NETCONF.
- Compare State: The module compares the running configuration against the declared parameters (
192.168.1.100 prefer: true). - Calculate Delta: If the NTP server is already configured with identical parameters, the delta is empty. The module exits immediately, reporting
ok(green) with zero changes made. - Apply Delta: If the NTP server is missing or configured without
prefer, the module generates and sends only the specific commands needed to achieve the target state, reportingchanged(yellow).
Practical Enterprise Scenario: Ansible Playbook Walkthrough
The following production-grade Ansible playbook demonstrates how network teams deploy NTP, domain name, and VLAN configurations to Cisco IOS-XE switches:
---
- name: Harden and Configure Campus Access Switches
hosts: campus_switches
gather_facts: false
connection: network_cli
vars:
primary_ntp: 192.168.10.50
secondary_ntp: 192.168.10.51
dns_domain: enterprise.lan
tasks:
- name: Configure System Domain Name and DNS Resolvers
cisco.ios.ios_system:
domain_name: "{{ dns_domain }}"
name_servers:
- 192.168.1.10
- 192.168.1.11
state: present
- name: Ensure Corporate VLANs Exist with Consistent Naming
cisco.ios.ios_vlans:
config:
- vlan_id: 10
name: MANAGEMENT
- vlan_id: 20
name: DATA_USERS
- vlan_id: 30
name: VOICE_VOIP
state: merged
- name: Configure Global NTP Infrastructure
cisco.ios.ios_ntp_global:
config:
servers:
- server: "{{ primary_ntp }}"
prefer: true
- server: "{{ secondary_ntp }}"
prefer: false
state: merged
- name: Save Running Configuration to Startup Configuration
cisco.ios.ios_config:
save_when: modified
Playbook Architectural Analysis
connection: network_cli: Informs Ansible to use specialized network plugins that maintain persistent SSH connections tailored for Cisco CLI command modes, rather than attempting to copy and execute Python files on the target switch.gather_facts: false: Prevents Ansible from running Linux-specific fact-gathering commands (likesetup.py), which would fail on standard switch CLIs.state: merged: Merges the declared configuration with the running configuration without deleting other unmentioned VLANs or NTP servers. Alternatively,state: replacedoroverriddenallows strict pruning of unmanaged configurations.save_when: modified: Enforces idempotency during configuration persistence. The switch issuescopy running-config startup-configonly if preceding tasks modified the running configuration, avoiding unnecessary flash wear.
Why is Ansible widely adopted for enterprise network device automation compared to traditional agent-based platforms like Chef and Puppet?
Ansible requires compiling C++ binaries for every switch architecture before execution
Ansible runs directly in the kernel space of Cisco switches, using proprietary microcode that Cisco loads onto each device
Ansible is agentless, using SSH, NETCONF, or APIs, so nothing has to be installed on closed network operating systems
Ansible utilizes pure Ruby cookbooks that run inside the switch bootloader
An automation engineer runs an idempotent Ansible playbook to configure an existing VLAN on a core switch. If the VLAN and its name already match the declared state in the playbook, what will Ansible report upon execution?
failed (red), because duplicate resource declarations trigger a syntax exception
ok (green), because the current state already matches the desired state, resulting in zero modifications
changed (yellow), because the task re-sends the CLI commands to overwrite the existing configuration
skipped (blue), because network tasks cannot inspect existing switch configurations
A network engineering team must execute a coordinated, simultaneous routing protocol migration across two redundant core switches. Which architectural communication model is best suited for this task and why?
A push model, because the control node dispatches changes immediately on demand, ensuring synchronized multi-device cutover
A pull model, because target switches query a central master at random intervals to avoid network congestion
A peer-to-peer broadcast model, because switches exchange configuration manifests using OSPF LSAs
A manual CLI copy-paste model, because automated configuration management cannot support routing protocols
Sections you finish are checked off in the contents.