{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "74278f9f",
   "metadata": {},
   "source": [
    "# Temporal and panel twins\n",
    "\n",
    "Time series and multi-environment (panel) data get their own twin kinds with their own machinery:\n",
    "lagged dependencies, latent influence detection, per-environment models, forecasts with\n",
    "attribution, and interventions that are scheduled in time. This notebook runs all of it live.\n",
    "\n",
    "| Kind | Kwargs | What it models |\n",
    "| --- | --- | --- |\n",
    "| `temporal` | `time=` | one time series with lagged causal structure |\n",
    "| `multi-environment-static` | `entity=` | the same system observed across environments |\n",
    "| `multi-environment-temporal` | `time=` + `entity=` | a panel: many environments, each a time series |"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8ed563dc",
   "metadata": {},
   "outputs": [],
   "source": [
    "import os\n",
    "import numpy as np\n",
    "import pandas as pd\n",
    "\n",
    "import rootcause as rc\n",
    "\n",
    "rc.login(base_url=os.environ.get(\"ROOTCAUSE_BASE_URL\", \"https://platform.rootcause.ai\"))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "091ebf6b",
   "metadata": {},
   "source": [
    "## A panel: three stores, thirty months\n",
    "\n",
    "Ground truth: price suppresses demand, demand carries momentum (a lag), seasonality moves both,\n",
    "and each store runs at its own scale. Long format: one row per store per month."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "15916579",
   "metadata": {},
   "outputs": [],
   "source": [
    "rng = np.random.default_rng(11)\n",
    "months = pd.date_range(\"2024-01-01\", periods=30, freq=\"MS\")\n",
    "stores = {\"london\": 1.0, \"paris\": 0.8, \"berlin\": 1.25}\n",
    "rows = []\n",
    "for store, scale in stores.items():\n",
    "    demand_prev = 100.0 * scale\n",
    "    for i, month in enumerate(months):\n",
    "        season = 12 * np.sin(2 * np.pi * (i % 12) / 12)\n",
    "        price = 20 + 2 * np.sin(2 * np.pi * (i % 12) / 12 + 1) + rng.normal(0, 0.5)\n",
    "        demand = 0.55 * demand_prev + 60 * scale - 2.4 * price + season + rng.normal(0, 4)\n",
    "        revenue = price * demand * 0.1 + rng.normal(0, 3)\n",
    "        rows.append({\"month\": month.strftime(\"%Y-%m-%d\"), \"store\": store,\n",
    "                     \"price\": round(price, 2), \"demand\": round(demand, 1), \"revenue\": round(revenue, 1)})\n",
    "        demand_prev = demand\n",
    "panel = pd.DataFrame(rows)\n",
    "panel.head()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "0b765f93",
   "metadata": {},
   "source": [
    "## Discovery finds more than edges\n",
    "\n",
    "`time=` and `entity=` make this a panel-temporal twin. Note the graph: alongside the causal\n",
    "edges, discovery surfaced a **latent influence**, a hidden common cause it detected in the data\n",
    "but could not name, and it attributes the store-level differences to the environment itself."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "9a70d93a",
   "metadata": {},
   "outputs": [],
   "source": [
    "graph = rc.discover(panel, time=\"month\", entity=\"store\", force=True)\n",
    "graph.edges"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "d326e7b2",
   "metadata": {},
   "outputs": [],
   "source": [
    "twin = graph.train()\n",
    "twin"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "55afa21d",
   "metadata": {},
   "source": [
    "## Per-environment sampling\n",
    "\n",
    "Panel twins hold one model per environment. Sampling narrows with `environments=` and the\n",
    "returned frame carries an `environment` column; seeds derive stable per-environment children,\n",
    "so comparisons are deterministic."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "fc824487",
   "metadata": {},
   "outputs": [],
   "source": [
    "draws = twin.sample(n=500, environments=[\"london\", \"berlin\"], seed=3)\n",
    "draws.to_frame().groupby(\"environment\").mean(numeric_only=True).round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4a97e57e",
   "metadata": {},
   "source": [
    "## Forecasts carry their reasoning\n",
    "\n",
    "Each forecast step comes with bounds and an attribution: how much of the prediction is trend,\n",
    "season, and each causal parent (with lags). `aggregate=\"sum\"` adds a combined series across\n",
    "environments; `origin_timestamp` anchors backtests."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1f2fe754",
   "metadata": {},
   "outputs": [],
   "source": [
    "fc = twin.forecast(horizon=6, targets=[\"revenue\"], aggregate=\"sum\")\n",
    "fc.to_frame()[[\"environment\", \"timestamp\", \"prediction\", \"lowerBound\", \"upperBound\"]].head(8).round(1)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c1a2c6fd",
   "metadata": {},
   "outputs": [],
   "source": [
    "fc.to_frame().loc[0, \"attribution\"]"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1c677440",
   "metadata": {},
   "source": [
    "## Interventions scheduled in time\n",
    "\n",
    "`rc.at` wraps any intervention value with when it applies: `persistent=True` from the first\n",
    "step onwards, `duration_steps=` for a limited window, `timestamp=` for a specific start.\n",
    "Here: a permanent 10 percent price cut, in London only."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "4ebf3cc7",
   "metadata": {},
   "outputs": [],
   "source": [
    "result = twin.intervene(\n",
    "    {\"price\": rc.at(rc.pct(-10), persistent=True)},\n",
    "    outcomes=[\"revenue\"],\n",
    "    environments=[\"london\"],\n",
    ")\n",
    "result"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Working with a subset of environments\n",
    "\n",
    "`twin.environments` lists what the panel actually holds — one row per\n",
    "environment with its sample size:"
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "twin.environments"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "`twin.env(...)` pins a handle to some of them. Its `graph` re-aggregates the\n",
    "causal adjacency over just those environments — edges carry `agreementRate`,\n",
    "the share of the subset's environments in which discovery found the\n",
    "relationship — and `adjacency(agreement_threshold=...)` turns the edge-survival\n",
    "knob (default 0.5). Every simulation on the handle is scoped automatically."
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "eu = twin.env(\"london\", \"berlin\")\n",
    "adjacency = eu.graph\n",
    "print(f\"{adjacency.attrs['envCount']} of {adjacency.attrs['totalEnvCount']} environments, \"\n",
    "      f\"threshold {adjacency.attrs['agreementThreshold']}\")\n",
    "adjacency[[\"source\", \"target\", \"strength\", \"agreementRate\"]]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Simulations on the handle run only in the subset — same verbs, pre-scoped.\n",
    "The intervention below asks what a price cut does to revenue in London and\n",
    "Berlin, leaving Paris untouched:"
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "eu.intervene({\"price\": rc.at(rc.pct(-10), persistent=True)}, outcomes=[\"revenue\"])"
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "eu.forecast(horizon=3, targets=[\"revenue\"]).to_frame()[\n",
    "    [\"environment\", \"timestamp\", \"prediction\", \"lowerBound\", \"upperBound\"]\n",
    "].round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Environments can also be selected **by their data** instead of by name — `where=`\n",
    "filters on any twin column through per-environment statistics, so \"the\n",
    "high-revenue stores\" needs no hand-maintained list. A tuple is\n",
    "`(column, op, value)` for constant-per-environment columns, or\n",
    "`(column, reduce, op, value)` with reduce one of `avg`, `min`, `max`, or `any`;\n",
    "`.environments` shows exactly what matched before you run anything:"
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "high_rev = twin.env(where=[(\"revenue\", \"avg\", \">\", 100)])\n",
    "high_rev.environments"
   ]
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": [
    "high_rev.forecast(horizon=2, targets=[\"revenue\"]).to_frame()[\n",
    "    [\"environment\", \"timestamp\", \"prediction\"]\n",
    "].round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e6e055ab",
   "metadata": {},
   "source": [
    "## Monthly refresh: assimilate instead of retrain\n",
    "\n",
    "When next month's rows arrive, the model doesn't need rebuilding — extend the twin's\n",
    "source and fold the new rows into the fitted model with `update()`. It finishes with a\n",
    "status, never an error: `committed` (rows folded in), `up_to_date` (nothing new), or\n",
    "`retrain_required` (the model can't take these rows incrementally — call `twin.retrain()`).\n",
    "\n",
    "Static and temporal twins assimilate out of the box. Panel twins need the v2 panel\n",
    "engine (an opt-in in the twin builder); on anything else `update()` simply reports\n",
    "`retrain_required`, and `twin.update_eligibility` tells you in advance. London's series\n",
    "as its own temporal twin:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a9af8d74",
   "metadata": {},
   "outputs": [],
   "source": [
    "london = panel[panel[\"store\"] == \"london\"][[\"month\", \"price\", \"demand\", \"revenue\"]].reset_index(drop=True)\n",
    "monthly = rc.discover(london, time=\"month\", force=True).train()\n",
    "monthly"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "65148237",
   "metadata": {},
   "source": [
    "Two months pass. Extend the source with the new rows and update — the transcript is\n",
    "the whole loop:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b67b8bca",
   "metadata": {},
   "outputs": [],
   "source": [
    "source = monthly.source\n",
    "current = source.to_frame()\n",
    "last_month = pd.to_datetime(current[\"month\"]).max()\n",
    "demand_prev = current.sort_values(\"month\")[\"demand\"].iloc[-1]\n",
    "\n",
    "rows = []\n",
    "for month in pd.date_range(last_month + pd.offsets.MonthBegin(1), periods=2, freq=\"MS\"):\n",
    "    j = (month.year - 2024) * 12 + (month.month - 1)\n",
    "    season = 12 * np.sin(2 * np.pi * (j % 12) / 12)\n",
    "    price = 20 + 2 * np.sin(2 * np.pi * (j % 12) / 12 + 1) + rng.normal(0, 0.5)\n",
    "    demand = 0.55 * demand_prev + 60 - 2.4 * price + season + rng.normal(0, 4)\n",
    "    rows.append({\"month\": month.strftime(\"%Y-%m-%d\"),\n",
    "                 \"price\": round(price, 2), \"demand\": round(demand, 1),\n",
    "                 \"revenue\": round(price * demand * 0.1 + rng.normal(0, 3), 1)})\n",
    "    demand_prev = demand\n",
    "\n",
    "source.extend(pd.DataFrame(rows))\n",
    "result = monthly.update()\n",
    "result"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b3f42fc4",
   "metadata": {},
   "source": [
    "Running it again with nothing new in the source is how a scheduled job stays honest —\n",
    "the second call is a cheap no-op:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b94dc57f",
   "metadata": {},
   "outputs": [],
   "source": [
    "monthly.update()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "bffcf5f2",
   "metadata": {},
   "source": [
    "The refreshed model forecasts onwards from the assimilated months:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "4d910beb",
   "metadata": {},
   "outputs": [],
   "source": [
    "fc = monthly.forecast(horizon=3, targets=[\"revenue\"])\n",
    "fc.to_frame()[[\"timestamp\", \"prediction\", \"lowerBound\", \"upperBound\"]].round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "c53786b0",
   "metadata": {},
   "source": [
    "The same scheduling works on plain temporal twins, and everything else from the\n",
    "[quickstart](quickstart.ipynb) applies unchanged: `save`, `load_twin`, `console`, and the\n",
    "ontology all understand temporal and panel twins."
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3 (ipykernel)",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.14.4"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}