Part 0

Foundations

This guide teaches you how to author DAGs (workflows) for Maestro-Pi — the Python syntax, the parameters each feature accepts, and how each feature actually behaves once your DAG is running.

What is a DAG?#

A DAG (Directed Acyclic Graph) is just Maestro-Pi's word for "a workflow." You write it as a plain Python file. It describes:

  • What work needs to happen — a set of individual steps, called tasks.
  • In what order — which tasks must finish before others can start.
  • When it should run — on a schedule, on demand, or in response to an event.

"Directed" means dependencies point one way (task A leads to task B, not the reverse). "Acyclic" means there are no loops — a task can never depend on itself, even indirectly through other tasks.

The smallest possible DAG#

python
from datetime import datetime
from dag_parser.dynamic.dag_context import DAG
from dag_parser.dynamic.operators import PythonOperator

with DAG(
    dag_id="hello_world",
    schedule="@daily",
    start_date=datetime(2026, 1, 1),
) as dag:

    def say_hello():
        print("Hello from PI-Flow!")

    hello_task = PythonOperator(
        task_id="say_hello",
        python_callable=say_hello,
    )

Save this as a .py file in the dags/ folder of your team's Git repo. Maestro-Pi periodically pulls that repo, notices the new/changed file, and parses it — no manual upload step. From that point on, hello_world shows up in the UI and starts producing scheduled runs.

The building blocks, in plain terms#

TermWhat it means
DAGThe whole workflow definition — one Python file (or one with DAG(...) block) usually equals one DAG.
TaskOne unit of work inside a DAG — "run this Python function," "run this SQL," "call this API."
OperatorThe kind of task — PythonOperator, BashOperator, SnowflakeOperator, etc. Each operator knows how to execute one kind of work.
DependencyAn ordering rule between two tasks, written with >> — "B only starts after A."
DAG RunOne execution of the whole DAG — e.g. "today's 2 AM run." A DAG can have many runs over time, one per schedule tick (or per manual trigger).
Task InstanceOne task, inside one specific DAG run — e.g. "the say_hello task, from today's 2 AM run." This is the thing that actually has a state (running, success, failed, ...).

A useful mental model: the DAG is the blueprint; a DAG Run is one building built from that blueprint; a Task Instance is one room in that building.

The task state lifecycle#

Every task instance moves through states as its run progresses. You'll see these constantly in the UI and in this guide:

text
none → scheduled → queued → running → success
                                    ↘ failed
                                    ↘ up_for_retry → (back to scheduled)
                                    ↘ skipped
                                    ↘ deferred → (back to scheduled, once its wait condition is met)
                                    ↘ up_for_reschedule → (back to scheduled, for reschedule-mode sensors)
  • none — not ready yet; still waiting on upstream tasks or a gating condition.
  • scheduled / queued — ready to run; waiting for a worker to pick it up.
  • running — actively executing.
  • success / failed / skipped — a terminal outcome for this attempt.
  • up_for_retry — failed, but will automatically try again.
  • deferred — waiting on an async condition (a timer, an HTTP check) without holding a worker slot.
  • up_for_reschedule — a "reschedule-mode" sensor's way of waiting without holding a worker slot between checks.

How your DAG file actually gets picked up#

You don't need to know the internals to author DAGs, but the short version helps explain some behavior later in this guide: your DAG file lives in a Git repository; Maestro-Pi periodically fetches that repo, detects new/changed .py files, and parses them. Once parsed, your DAG's structure (tasks, schedule, dependencies) is recorded, and from then on Maestro-Pi's scheduler takes over — creating runs, planning tasks, and dispatching them to run, all without your file being touched again until you edit it and push a new commit.

One thing worth knowing early: an in-flight run is never affected by a later edit to the DAG file. Every run is pinned to the exact version of the DAG that existed when it started, so you can safely fix bugs or add tasks without disturbing runs already in progress.

How this guide is organized#

  1. Declaring a DAG & Scheduling — the basics: identity, when it runs, time window, catchup.
  2. Controlling a DAG Run's behavior — concurrency, timeouts, SLAs, typed parameters, callbacks, access control, partitioning, templating options.
  3. Tasks — behavior & customization — retries, trigger rules, cross-run dependencies, timeouts, task SLA, pools, priority, setup/teardown, per-task callbacks, environment control.
  4. Scheduling & triggering runs — automatic runs, manual triggers, backfill, dataset triggers, cross-DAG triggers, deferred waits.
  5. Dependencies & flow control — edges, labels, convergence, dynamic mapping, branching, voluntary skip.
  6. Data passing & templating — XCom, TaskFlow, Go/Jinja templating, Variables & Connections.
  7. Alerting & notifications — email, Slack, HTTP webhook, PagerDuty, and how event scope works.
  8. Sensors & waiting — the four built-in sensors, tuning, poke vs. reschedule, deferrable waits.
  9. Reliability, in one place — a short recap of how SLA/timeout/retry mechanisms fit together.
  10. Operators & integrations catalog — the full list of what a task can actually do.

Each feature below follows the same four-part layout:

  • What it does — plain-language explanation.
  • Parameters — every knob, its type, its default, and its valid range/values.
  • How it works — the behavior you should understand before relying on it (not a bug list — just how the mechanism actually operates).
  • Example — working code.