I Turned a Three-Hour Network Maintenance Check Into Three Minutes With Python
It was 3:00 AM on a Tuesday, and I was staring at three terminal windows while my coffee slowly turn 2026-9-28 15:38:28 Author: hackernoon.com(查看原文) 阅读量:3 收藏

It was 3:00 AM on a Tuesday, and I was staring at three terminal windows while my coffee slowly turned cold. Earlier that night, our team had updated dozens of network switches. Now came the part nobody enjoys: running the same health checks on every device, copying the output, and comparing it with the state we had recorded before the change.

If you work in infrastructure, you know how this goes. The commands are not especially difficult, but the process becomes fragile when you repeat it hundreds of times. At that hour, a missed route or inactive interface can turn into a morning outage ticket.

When the scope grew to 200 devices, the manual approach finally stopped making sense. I built a small Python utility that connected to devices in parallel, ran a standard set of checks, saved the results, and separated failures for another attempt. The first version cut a task that took more than three hours down to about three minutes.

The Real Problem Was Repetition

At first, I thought I was automating a few command line steps. That description was technically true, but it missed the important part. I was really trying to remove hundreds of tiny decisions from a maintenance window.

A person working through a device list has to remember which host is next, which commands have already run, where each output file belongs, and whether a strange result is a command error or a connection failure. None of those decisions is hard on its own. Put them together at 3:00 AM, across 200 devices, and the process becomes an accuracy problem.

The safest automation did not need to be clever. It needed to make the boring path consistent and make exceptions obvious. That became the design goal for the whole tool.

What the Automation Actually Did

The workflow had four basic inputs and outputs. It read a list of target devices, loaded a standard set of health check commands, asked the operator for credentials, and created a dated folder for the results. From there, a pool of workers opened many Secure Shell connections at the same time and ran the same checks on each device.

Every device received its own log. That detail mattered more than it might sound, because one giant stream of mixed terminal output is almost impossible to review during an incident. A per-device record made it easy to compare results, send a specific file to another engineer, or revisit one questionable node without digging through everything else.

The script also produced two failure views. One contained a clean list of devices that failed, which could be used for a targeted rerun. The other stored the detailed error information needed for troubleshooting. The operator could act quickly without losing the context required for a proper investigation later.

Why Parallel Checks Changed the Math

Network checks spend a surprising amount of time waiting. A device has to accept the connection, authenticate the user, process a command, and send the response back. If a script checks devices one after another, most of its life is spent waiting for remote systems.

Running several connections concurrently changes that. While one device is thinking, the program can talk to another. The total time starts to depend less on the sum of every delay and more on the slowest groups of devices.

I used up to 100 worker threads because the work was dominated by network input and output. That number is not a universal recommendation. A good limit depends on device capacity, authentication services, network conditions, jump hosts, and operational policy, so it should be tested gradually in the real environment.

A Fast Script Still Needs Brakes

Concurrency can make a safe process faster, but it can also make a bad process fail faster. I kept the first use case deliberately narrow: read-only health checks before and after a planned change. The tool was collecting evidence, not making configuration decisions on its own.

Timeouts prevented one unreachable device from holding the entire run hostage. Each connection was closed after its work finished, even when a command failed. Credentials were entered at runtime instead of being stored in the script or a plain text file.

I also treated the device and command lists as change inputs that deserved review. Before starting a large run, the operator could confirm exactly which commands had loaded. Testing against a small batch first was much safer than discovering a syntax mismatch across the full fleet.

Failure Had to Stay Local

One early design choice made the tool far more useful in practice: a failed command did not end the entire session. A device might reject one command because its software version uses different syntax, yet still respond perfectly to the rest of the health checks. The script recorded the error and moved on.

The same principle applied at the device level. A bad password, timeout, or unreachable address affected that target only. Other workers kept going, and the failed device appeared in the rerun list.

This is the difference between a demo and an operations tool. A demo assumes the happy path. A tool used during maintenance has to assume that at least one device, command, or connection will behave strangely and still produce a useful result.

Logs Became the Product

The first thing people notice is the speed improvement, but the structured output was just as valuable. Manual copy and paste creates inconsistent evidence. File names drift, timestamps disappear, and two engineers may record the same check in different ways.

A dated directory and one log per device created a repeatable record of the maintenance window. I could see when the check ran, whether the session succeeded, which commands produced output, and which devices needed attention. That made handoffs and follow-up work much easier.

It also changed how I thought about the script. The connections and threads were implementation details. The actual product was a trustworthy set of records that helped a tired engineer decide whether the network was healthy enough to close the change.

The Check That Justified the Work

On the second run, the automated checks flagged an inactive interface on a core node after a firmware reload. It was exactly the kind of small line in a large output that is easy to miss when someone has been copying results for hours.

Because the issue appeared immediately in the device log, we could investigate before users started reporting packet loss. I cannot prove that every manual review would have missed it, but the automation shortened the time between the change and the discovery. That is the kind of advantage that matters during a maintenance window.

The tool also gave the team a clean stopping condition. Instead of relying on a vague feeling that everything looked fine, we had a completed device list, captured results, and a short exception list to resolve.

What I Would Improve Next

The simple version solved the immediate problem, but I would not treat it as the final form. The next useful step would be to compare selected pre-change and post-change values automatically, then highlight meaningful differences. Raw logs are good evidence, but a concise summary helps the operator focus.

I would also add controlled retry behavior, better support for different device types, and a clear concurrency setting for each environment. Metrics such as success rate, connection time, command time, and retry count would make performance problems easier to spot.

For a larger platform, I would move credentials into an approved secrets system and connect the results to the change record. Those improvements should come after the basic workflow is trusted. Adding features before the failure model is clear usually creates a more impressive tool, not a safer one.

The Lesson I Took Away

The biggest win did not come from an advanced algorithm. It came from looking at a repetitive operational task and asking which parts required judgment and which parts only required consistency. The computer handled the connections, commands, timestamps, and files whereas I stayed responsible for interpreting exceptions and deciding whether to proceed.

That division of work turned a three-hour manual check into a roughly three-minute automated run. More importantly, it made the result easier to trust. At 3:00 AM, that is worth far more than a clever code sample.

from datetime import datetime
from getpass import getpass
import logging

from pathlib import Path
from concurrent.futures import ThreadPoolExecutor, as_completed
from netmiko import ConnectHandler

# Paths & Dynamic Timestamps
HOSTS_FILE = Path('hosts.txt')
PRECHECK_COMMANDS_FILE = Path('config.txt')

TIMESTAMP = datetime.now().strftime("%Y%m%d_%H%M%S")
RUN_DIR = Path('./logs') / f"prechecks_{TIMESTAMP}"
SCRIPT_LOGS_DIR = RUN_DIR / 'script_logs'
HOST_LOGS_DIR = RUN_DIR / 'host_output'
GLOBAL_LOG = RUN_DIR / 'global.log'

# Ensure isolated directory tree exists for this run
SCRIPT_LOGS_DIR.mkdir(parents=True, exist_ok=True)
HOST_LOGS_DIR.mkdir(parents=True, exist_ok=True)

# Set up logging for this run
logging.basicConfig(
    filename=GLOBAL_LOG,
    level=logging.DEBUG,
    format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger("global")


def write_failures_list(failed_host):
    failure_log = SCRIPT_LOGS_DIR / 'failures_list.log'
    with open(failure_log, 'a') as failure_file:
        failure_file.write(f"{failed_host}\n")
    logger.info(f"Logged failure for {failed_host} in failures_list.log")


def write_failures_extensive(failed_host_and_error_details):
    failure_log = SCRIPT_LOGS_DIR / 'failures_extensive.log'
    with open(failure_log, 'a') as failure_file:
        failure_file.write(f"{failed_host_and_error_details}\n")
    logger.error(f"Logged extensive failure: {failed_host_and_error_details}")


def run_pre_checks(host_ip, username, password, commands_to_run, logs_directory):
    host_dict = {
        'device_type': 'cisco_ios',
        'host': host_ip,
        'username': username,
        'password': password,
        'secret': password,
        'conn_timeout': 30,
    }
    
    precheck_output_accumulator = []
    host_encountered_errors = False

    try:
        logger.info(f"Attempting connection to {host_ip}")
        with ConnectHandler(**host_dict) as net_connect:
            net_connect.enable()
            
            for command in commands_to_run:
                command = command.strip()
                if not command:
                    continue
                
                # Isolated per-command execution 
                try:
                    output = net_connect.send_command(command)
                    precheck_output_accumulator.append(f"--- {command} ---\n{output}\n")
                except Exception as cmd_error:
                    host_encountered_errors = True
                    error_detail = f"ERROR running '{command}': {str(cmd_error)}"
                    logger.warning(f"[{host_ip}] {error_detail}")
                    precheck_output_accumulator.append(f"--- {command} ---\n[EXECUTION ERROR]: {cmd_error}\n")
                    write_failures_extensive(f"{host_ip} - Command '{command}' failed: {str(cmd_error)}")

        # Write per-device output log
        host_log_file = logs_directory / f"{host_ip}_precheck.log"
        with open(host_log_file, 'w') as log_file:
            log_file.write("\n".join(precheck_output_accumulator))

        if host_encountered_errors:
            # Device partially succeeded; log to failures list for review
            write_failures_list(host_ip)
            logger.warning(f"Finished pre-checks on {host_ip} with partial command errors.")
            return False

        logger.info(f"Successfully finished all pre-checks on {host_ip}")
        return True

    except Exception as e:
        # Host-level connection/auth failure
        error_msg = f"{host_ip}: Connection/Execution failed - {str(e)}"
        write_failures_list(host_ip)
        write_failures_extensive(error_msg)
        return False


def main():
    if not HOSTS_FILE.exists() or not PRECHECK_COMMANDS_FILE.exists():
        logger.error("Missing hosts.txt or config.txt file.")
        print("Error: Ensure 'hosts.txt' and 'config.txt' exist in the current directory.")
        return

    username = input("Username: ")
    password = getpass("Password: ")

    with open(HOSTS_FILE, 'r') as f:
        hosts = [line.strip() for line in f if line.strip()]

    with open(PRECHECK_COMMANDS_FILE, 'r') as f:
        commands = [line.strip() for line in f if line.strip()]
    
    max_threads = 100
    print(f"Starting pre-checks across {len(hosts)} hosts using up to {max_threads} threads...")
    print(f"Results will be saved to: {RUN_DIR}")

    with ThreadPoolExecutor(max_workers=max_threads) as executor:
        futures = {
            executor.submit(run_pre_checks, host, username, password, commands, HOST_LOGS_DIR): host
            for host in hosts
        }

        for future in as_completed(futures):
            host = futures[future]
            try:
                success = future.result()
                if success:
                    print(f"[+] Completed: {host}")
                else:
                    print(f"[-] Encountered Errors/Failed: {host}")
            except Exception as exc:
                logger.error(f"{host} generated an unhandled exception: {exc}")
                print(f"[!] Error processing {host}: {exc}")


if __name__ == "__main__":
    main()

文章来源: https://hackernoon.com/i-turned-a-three-hour-network-maintenance-check-into-three-minutes-with-python?source=rss
如有侵权请联系:admin#unsafe.sh