{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "745760b5",
   "metadata": {},
   "source": [
    "# RootCause SDK quickstart\n",
    "\n",
    "Causal discovery, digital twins, and what-if questions from Python. Two modes, one object model:\n",
    "\n",
    "- **Direct mode** — `rc.discover(df)` on a pandas DataFrame, zero workspace ceremony.\n",
    "- **Platform mode** — `rc.workspace(...)` over everything your team builds in the RootCause UI.\n",
    "\n",
    "```bash\n",
    "pip install rootcause-sdk\n",
    "```"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8eafe714",
   "metadata": {},
   "outputs": [],
   "source": [
    "import os\n",
    "import numpy as np\n",
    "import pandas as pd\n",
    "\n",
    "import rootcause as rc\n",
    "\n",
    "rc.login(base_url=os.environ.get(\"ROOTCAUSE_BASE_URL\", \"https://platform.rootcause.ai\"))\n",
    "rc.workspaces()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "ba1c3751",
   "metadata": {},
   "source": [
    "## Direct mode: from DataFrame to causal graph\n",
    "\n",
    "A synthetic marketing funnel with known ground truth: `marketing_spend -> leads -> revenue`,\n",
    "with `seasonality` nudging leads. Discovery should recover exactly that structure."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "d832c792",
   "metadata": {},
   "outputs": [],
   "source": [
    "rng = np.random.default_rng(7)\n",
    "n = 240\n",
    "marketing = rng.normal(50, 12, n)\n",
    "seasonality = rng.normal(0, 1, n)\n",
    "leads = 3.0 * marketing + 15 * seasonality + rng.normal(0, 8, n)\n",
    "revenue = 2.2 * leads + rng.normal(0, 20, n)\n",
    "\n",
    "df = pd.DataFrame({\n",
    "    \"marketing_spend\": marketing.round(2),\n",
    "    \"seasonality\": seasonality.round(3),\n",
    "    \"leads\": leads.round(1),\n",
    "    \"revenue\": revenue.round(1),\n",
    "})\n",
    "df.head()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "75cb5fe0",
   "metadata": {},
   "outputs": [],
   "source": [
    "graph = rc.discover(df)\n",
    "graph.edges"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "3c278f3d",
   "metadata": {},
   "source": [
    "The adjacency matrix is a labelled DataFrame; `.to_numpy()` and `.to_networkx()` are there when you need them.\n",
    "\n",
    "Re-running `rc.discover(df)` on identical data reuses this twin instantly. If a model is ever\n",
    "corrupt or predates an engine fix, `rc.discover(df, force=True)` re-runs discovery from scratch."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "f1d9edfc",
   "metadata": {},
   "outputs": [],
   "source": [
    "graph.adjacency()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "c04cdb0c",
   "metadata": {},
   "source": [
    "## Domain knowledge, then training\n",
    "\n",
    "`pin` fixes an edge as present; `forbid` as absent. Training fits the causal Bayesian network."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "7a26e937",
   "metadata": {},
   "outputs": [],
   "source": [
    "graph.pin(\"marketing_spend\", \"leads\")\n",
    "twin = graph.train()\n",
    "twin"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1618464c",
   "metadata": {},
   "source": [
    "## The power-user primitive: raw joint draws\n",
    "\n",
    "Every simulation family wraps conditional sampling. `twin.sample()` hands you the draws so you can\n",
    "compute your own estimands. Seeds are reproducible across every twin family."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8a17863f",
   "metadata": {},
   "outputs": [],
   "source": [
    "draws = twin.sample(n=2000, seed=42)\n",
    "draws.to_frame().describe().round(1)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "f876d09d",
   "metadata": {},
   "outputs": [],
   "source": [
    "boosted = twin.sample(n=2000, do={\"marketing_spend\": rc.pct(+20)}, seed=42)\n",
    "pd.DataFrame({\n",
    "    \"baseline\": draws.to_frame().mean(),\n",
    "    \"do(marketing +20%)\": boosted.to_frame().mean(),\n",
    "}).round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "68bd3ec1",
   "metadata": {},
   "source": [
    "## Interventions, narrated\n",
    "\n",
    "`intervene` runs the full simulation machinery server-side and blocks for the result."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "54d5d335",
   "metadata": {},
   "outputs": [],
   "source": [
    "result = twin.intervene({\"marketing_spend\": rc.pct(+25)}, outcomes=[\"revenue\", \"leads\"])\n",
    "result"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6d5f7dba",
   "metadata": {},
   "source": [
    "## Ontology queries\n",
    "\n",
    "Every upload gets ontology concepts. The query engine handles filters, aggregations, grouping,\n",
    "and natural language — `select` takes concept names."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "cbfe7b42",
   "metadata": {},
   "outputs": [],
   "source": [
    "from rootcause.direct import scratch_workspace\n",
    "onto = scratch_workspace(rc._transport()).ontology\n",
    "onto.concepts"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a1ccbc7d",
   "metadata": {},
   "outputs": [],
   "source": [
    "onto.query(select=[\"Revenue\", \"Leads\"], limit=5).to_frame(max_rows=5)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "64837547",
   "metadata": {},
   "source": [
    "## Interactive apps under the cell\n",
    "\n",
    "The same interactive consoles Claude and ChatGPT render for RootCause tools work as notebook\n",
    "widgets: the twin console below explores the graph and re-runs scenarios live through the\n",
    "platform. Requires `pip install \"rootcause-sdk[jupyter]\"`; the widget renders in JupyterLab,\n",
    "Notebook 7, VS Code, and Colab (static exports show a placeholder)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "599004df",
   "metadata": {},
   "outputs": [],
   "source": [
    "twin.console(height=560)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "11926995",
   "metadata": {},
   "source": [
    "## Portable twins\n",
    "\n",
    "The export zip carries the trained model parameters; `rc.load_twin` brings it back anywhere."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "d5c05c2a",
   "metadata": {},
   "outputs": [],
   "source": [
    "path = twin.save(\"quickstart.rctwin\")\n",
    "f\"{path.name}: {path.stat().st_size:,} bytes\""
   ]
  },
  {
   "cell_type": "markdown",
   "id": "23ee64d5",
   "metadata": {},
   "source": [
    "## Platform mode\n",
    "\n",
    "The same classes over shared workspaces — twins your colleagues trained in the UI are just there:\n",
    "\n",
    "```python\n",
    "ws = rc.workspace(\"Calix Forecasting\")\n",
    "ws.sources[\"shipments\"].to_frame()\n",
    "\n",
    "twin = ws.twin(\"C8 Temporal\")\n",
    "fc = twin.forecast(horizon=24)\n",
    "fc.to_frame()\n",
    "\n",
    "twin.ask(\"what happens to bookings if we cut trade shows entirely?\")\n",
    "```"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3 (ipykernel)",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.14.4"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}