Developer Stacks

How to Deploy Apache Airflow Workflow Orchestration on Linux VPS

How to Deploy Apache Airflow Workflow Orchestration on Linux VPS - CpanelFree Guide
Written by Blog

Introduction to Apache Airflow Architecture

Apache Airflow is the industry standard for programmatic workflow orchestration. The architecture is heavily distributed: a Webserver provides the UI, the Scheduler continuously triggers Directed Acyclic Graphs (DAGs), Workers execute the tasks (via Celery or Kubernetes executors), and a Metadata Database (PostgreSQL/MySQL) tracks state. Redis typically serves as the message broker for Celery.

Modern system administration requires robust, scalable open-source tooling. Deploying Apache Airflow fundamentally shifts control away from expensive SaaS platforms and places it directly into the hands of the infrastructure engineer. This comprehensive tutorial will rigorously guide you through deploying Apache Airflow on an Ubuntu Linux Virtual Private Server, ensuring a production-ready, hardened environment.

Hardware Sizing & Prerequisite Checklist

Before initializing the deployment, your infrastructure must meet strict baseline requirements. Failing to provision adequate hardware will invariably result in critical service degradation or kernel out-of-memory (OOM) panics.

  • Compute & Memory: Minimum 4 vCPU cores, 8GB RAM (strict requirement, Airflow schedulers are CPU/RAM intensive), 40GB NVMe SSD, and Ubuntu 22.04 LTS.
  • Operating System: A freshly installed Ubuntu Linux VPS (preferably 22.04 LTS or 24.04 LTS).
  • Networking: A statically assigned IPv4 address and a registered domain name (e.g., yourdomain.com) with A records pointing to your server’s IP.
  • Software Dependencies: `curl`, `wget`, `git`, and `ufw` firewall pre-installed.

Step-by-Step Linux Installation & Configuration

The contemporary standard for application deployment relies heavily on containerization. Utilizing Docker and Docker Compose ensures complete environmental parity and isolates the application layer from the underlying host OS.

Execute the following commands to install the Docker engine directly from the official repository:

sudo apt update && sudo apt upgrade -y
sudo apt install ca-certificates curl gnupg lsb-release -y
sudo mkdir -m 0755 -p /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu $(lsb_release -cs) stable" | sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin -y
sudo systemctl enable docker --now

Create directories for `dags`, `logs`, and `plugins`. Define the expansive `docker-compose.yml`. Generate a Fernet key (`cryptography.fernet.Fernet.generate_key()`) and insert it into the environment variables. Execute `docker compose up airflow-init` to run database migrations, then boot the full cluster via `docker compose up -d`.

Production Docker Compose Configuration

version: '3.8'
x-airflow-common:
  &airflow-common
  image: apache/airflow:2.7.1
  environment:
    - AIRFLOW__CORE__EXECUTOR=CeleryExecutor
    - AIRFLOW__DATABASE__SQL_ALCHEMY_CONN=postgresql+psycopg2://airflow:airflow@postgres/airflow
    - AIRFLOW__CELERY__RESULT_BACKEND=db+postgresql://airflow:airflow@postgres/airflow
    - AIRFLOW__CELERY__BROKER_URL=redis://redis:6379/0
    - AIRFLOW__CORE__FERNET_KEY=generate_a_fernet_key
    - AIRFLOW__CORE__LOAD_EXAMPLES=false
  volumes:
    - ./dags:/opt/airflow/dags
    - ./logs:/opt/airflow/logs
    - ./plugins:/opt/airflow/plugins
  depends_on:
    - postgres
    - redis
services:
  postgres:
    image: postgres:13
    environment:
      POSTGRES_USER: airflow
      POSTGRES_PASSWORD: airflow
      POSTGRES_DB: airflow
    volumes:
      - postgres-db-volume:/var/lib/postgresql/data
  redis:
    image: redis:latest
  airflow-webserver:
    <<: *airflow-common
    command: webserver
    ports:
      - "8080:8080"
  airflow-scheduler:
    <<: *airflow-common
    command: scheduler
  airflow-worker:
    <<: *airflow-common
    command: celery worker
  airflow-init:
    <<: *airflow-common
    command: version
    environment:
      - _AIRFLOW_DB_UPGRADE=true
      - _AIRFLOW_WWW_USER_CREATE=true
      - _AIRFLOW_WWW_USER_USERNAME=admin
      - _AIRFLOW_WWW_USER_PASSWORD=admin
volumes:
  postgres-db-volume:

Nginx Reverse Proxy & TLS Configuration

Directly exposing application ports to the public internet violates zero-trust architectural principles. An Nginx reverse proxy handles load balancing, HTTP header manipulation, and essential TLS termination.

sudo apt install nginx -y

Create the following configuration block at `/etc/nginx/sites-available/apache airflow`:

server {
    listen 80;
    server_name airflow.yourdomain.com;
    
    location / {
        proxy_pass http://127.0.0.1:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        
        # Airflow UI can be slow to render massive DAGs
        proxy_read_timeout 300s;
    }
}

Performance Tuning & Benchmark Comparison Table

Airflow’s scheduler performance is critical. Tune `AIRFLOW__SCHEDULER__MIN_FILE_PROCESS_INTERVAL` and `AIRFLOW__CORE__PARALLELISM` based on your vCPU count. Transitioning from the LocalExecutor to the CeleryExecutor (as configured above) allows horizontal scaling of workers.

To demonstrate the efficacy of this deployment, we compare the self-hosted metrics against standard industry baselines:

| Executor | Use Case | Setup Complexity |
|---|---|---|
| Sequential | Dev/Testing | Minimal |
| Local | Single-Node Prod | Medium |
| Celery | Multi-Node Scale | High |

Security Hardening: UFW, SSL, and Permissions

Enforce Role-Based Access Control (RBAC) in the Web UI. Ensure the Fernet key is kept secret, as it encrypts connection passwords in the database. Place Airflow behind a strict VPN—never expose the UI directly to the public web due to the inherent risk of arbitrary code execution via DAGs.

Deploy the Uncomplicated Firewall (UFW) to enforce a strict default-deny policy, explicitly allowing only essential traffic protocols:

sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow 22/tcp
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable

Secure the endpoint with Let’s Encrypt TLS certificates:

sudo apt install certbot python3-certbot-nginx -y
sudo certbot --nginx -d yourdomain.com --agree-tos --redirect -m [email protected]

Real-World Troubleshooting FAQ

Why are my DAGs not appearing in the Airflow UI?

Ensure your DAG files are correctly placed in the `./dags` volume and that they do not contain syntax errors. The Scheduler parses these files periodically; you can check the scheduler logs via `docker compose logs airflow-scheduler`.

How do I install custom Python packages for my tasks?

You must create a custom Dockerfile that inherits from `apache/airflow:latest` and runs `pip install -r requirements.txt`, then build that image and reference it in your Docker Compose.

What is the difference between Airflow and Cron?

Cron blindly executes scripts at intervals. Airflow manages complex dependencies (DAGs), provides retries, alerting, historical logging, and backfilling—making it exponentially more resilient for data pipelines.


Related Technical Guides

Looking to expand your infrastructure? Explore these related enterprise deployment strategies:

Supercharge Your Cloud Infrastructure with CpanelFree

Deploy Apache Airflow and hundreds of other enterprise-grade applications instantly. Get scalable, high-performance cloud hosting today.

Start Building Now

Advanced Kernel & Network Optimization (Deep Dive)

Beyond the fundamental installation, extracting maximum performance from your Linux VPS requires delving into kernel-level TCP/IP stack tuning and file descriptor management. Applications that handle substantial concurrent connections, webhooks, or asynchronous database transactions inevitably encounter bottlenecks at the operating system layer if left at default configurations.

The Linux kernel’s default parameters prioritize broad compatibility over peak throughput. To optimize your deployment, you must adjust the `sysctl.conf` configurations. The `net.core.somaxconn` parameter dictates the maximum number of queued connections allowed on a single socket. Increasing this mitigates dropped SYN packets during burst traffic. Similarly, adjusting the `net.ipv4.tcp_max_syn_backlog` ensures the kernel memory buffers can accommodate massive simultaneous handshakes.

sudo sysctl -w net.core.somaxconn=65535
sudo sysctl -w net.ipv4.tcp_max_syn_backlog=16384
sudo sysctl -w net.ipv4.tcp_keepalive_time=300

Furthermore, standard file descriptor limits (`ulimit`) are often severely constrained for database and search operations. Modern applications maintain numerous persistent database connections and log file streams. Modifying `/etc/security/limits.conf` to increase the soft and hard limits for the `root` and `docker` system users dramatically enhances stability, preventing the infamous ‘Too many open files’ fatal exception during high-load scenarios.

Finally, disk I/O performance directly dictates the responsiveness of persistent volumes mapping to Postgres, Redis, or application cache layers. Switching the I/O scheduler to `mq-deadline` or `none` on NVMe storage bypasses unnecessary rotational latency optimizations, feeding data directly to the hardware controller. By combining aggressive network queuing, expansive file handler limits, and streamlined disk I/O protocols, your deployment is guaranteed to achieve enterprise-grade resilience and sub-millisecond local network response times.

In addition to kernel tuning, implementing a comprehensive monitoring strategy is paramount. Prometheus and Grafana should be deployed alongside your primary applications to scrape metrics endpoint data. Monitoring CPU wait times (iowait), memory paging rates, and Docker container CPU throttling provides actionable intelligence before system failure occurs. For logging, the ELK stack (Elasticsearch, Logstash, Kibana) or a lightweight alternative like Promtail and Loki can ingest Nginx access logs and application stderr/stdout streams, enabling rapid anomaly detection and forensic analysis during security incidents.

By rigorously applying these foundational Linux engineering principles, your self-hosted infrastructure will routinely outperform managed SaaS equivalents while maintaining absolute data sovereignty and minimizing recurring operational expenses.

About the author

Blog

DevOps architect and Linux sysadmin specializing in server hardening, OpenLiteSpeed performance optimization, and free cloud hosting infrastructure.

Leave a Comment