Home / Articles / Multi-Agent Debate for Volatility Breakout Optimization: Can Three Gemini

This article is published in English.

Multi-Agent Debate for Volatility Breakout Optimization: Can Three Gemini

Operable walkthrough of Multi-Agent Debate for Volatility Breakout Optimization: Can Three Gemini: contracts, checks, and drop-in code slots for teams shipping this pattern.

3043 words

the author following notes reconstruct a practical path around “Multi-Agent Debate for Volatility Breakout Optimization: Can Three Gemini Personas Fix a Losing Strategy?”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

{
    "strategy_version": "v2.0-optimized",
    "final_sharpe_ratio": -0.6822300207528283,
    "consensus_rationale": "Iteration 1 consensus achieved superior Sharpe ratio through optimized ATR multiplier and volume threshold refinement.",
    "max_drawdown_delta": 0.00687485272186061,
    "total_iterations_run": 3
}

Data Loading and Market Universe Setup

the author Data Loading and Market stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Designing the High-Beta Universe

the author Designing the High-Beta Universe stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

import os
import json
import logging
from datetime import datetime, timezone
import pandas as pd
import numpy as np
import yfinance as yf
import matplotlib.pyplot as plt
from google import genai
logging.basicConfig(level=logging.INFO,
                    format='%(asctime)s - %(levelname)s - %(message)s')

Handling Missing Data and MultiIndex Alignment

the author Handling Missing Data and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. the author Handling Missing Data and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

class MarketDataLoader:
    def __init__(self, tickers: list[str], benchmark: str):
        self.tickers = tickers
        self.all_symbols = tickers + [benchmark]
        self.benchmark = benchmark

    def fetch_data(self) -> tuple[pd.DataFrame, pd.DataFrame, pd.DataFrame, pd.DataFrame]:
        end_date = datetime.now(timezone.utc).strftime('%Y-%m-%d')
        start_date = (datetime.now(timezone.utc) -
                      pd.Timedelta(days=365 * 3)).strftime('%Y-%m-%d')
        logging.info(
            f"Fetching historical data from {start_date} to {end_date} for {len(self.all_symbols)} symbols.")

        data = yf.download(self.all_symbols, start=start_date,
                           end=end_date, auto_adjust=True, progress=False)
        if data.empty:
            logging.error(
                f"Failed to download data for symbols: {self.all_symbols}")
            raise ValueError(
                f"Failed to download data for symbols: {self.all_symbols}")

        if isinstance(data.columns, pd.MultiIndex):
            close = data['Close'] if 'Close' in data.columns.levels[0] else data.xs(
                'Close', level=1, axis=1)
            high = data['High'] if 'High' in data.columns.levels[0] else data.xs(
                'High', level=1, axis=1)
            low = data['Low'] if 'Low' in data.columns.levels[0] else data.xs(
                'Low', level=1, axis=1)
            volume = data['Volume'] if 'Volume' in data.columns.levels[0] else data.xs(
                'Volume', level=1, axis=1)
        else:
            close = data[['Close']]
            high = data[['High']]
            low = data[['Low']]
            volume = data[['Volume']]

        close = close.ffill().dropna(how='all')
        high = high.ffill().dropna(how='all')
        low = low.ffill().dropna(how='all')
        volume = volume.ffill().fillna(0.0)

        if close.empty:
            logging.error(
                "Cleaned price data DataFrame is empty after processing.")
            raise ValueError(
                "Cleaned price data DataFrame is empty after processing.")

        logging.info(
            f"Data checkpoint - Shape: {close.shape}, Column sample: {list(close.columns[:3])}, NaN count: {close.isna().sum().sum()}, Date range: {close.index[0]} to {close.index[-1]}")
        return close, high, low, volume
INFO - Fetching historical data from 2023-08-10 to 2026-08-09 for 16 symbols.
INFO - Data checkpoint - Shape: (751, 16), Column sample: ['AMD', 'COIN', 'DKNG'], NaN count: 0, Date range: 2023-08-10 00:00:00 to 2026-08-07 00:00:00

Defining the Volatility Breakout Engine

For the Defining the Volatility Breakout stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

the author Mechanics of Momentum and Volatility

For the the author Mechanics of Momentum stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Simulating Portfolio Returns

For the Simulating Portfolio Returns stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Simulating Portfolio Returns stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

class VolatilityStrategyBacktester:
    def __init__(self, close_df: pd.DataFrame, high_df: pd.DataFrame, low_df: pd.DataFrame, volume_df: pd.DataFrame, benchmark: str):
        self.close_df = close_df
        self.high_df = high_df
        self.low_df = low_df
        self.volume_df = volume_df
        self.benchmark = benchmark
        self.asset_columns = [c for c in close_df.columns if c != benchmark]

    def backtest(self, params: dict) -> tuple[pd.Series, pd.Series]:
        lookback = int(params.get("lookback", 20))
        atr_mult = float(params.get("atr_multiplier", 2.0))
        vol_thresh = float(params.get("volume_threshold", 1.5))

        asset_returns_dict = {}

        for ticker in self.asset_columns:
            if ticker not in self.close_df.columns:
                continue
            c = self.close_df[ticker]
            h = self.high_df[ticker]
            l = self.low_df[ticker]
            v = self.volume_df[ticker]

            tr1 = h - l
            tr2 = (h - c.shift(1)).abs()
            tr3 = (l - c.shift(1)).abs()
            tr = pd.concat([tr1, tr2, tr3], axis=1).max(axis=1)
            atr = tr.rolling(window=lookback).mean()

            rolling_high = c.rolling(window=lookback).max()
            avg_vol = v.rolling(window=lookback).mean()

            breakout = (c > rolling_high.shift(1) + atr_mult *
                        atr) & (v > vol_thresh * avg_vol)

            asset_ret = c.pct_change()
            strat_ret = asset_ret * breakout.shift(1).fillna(False)

            trades = breakout & (~breakout.shift(1).fillna(False))
            strat_ret = strat_ret - (trades.astype(float) * 0.001)

            asset_returns_dict[ticker] = strat_ret

        strategy_df = pd.DataFrame(asset_returns_dict)
        portfolio_returns = strategy_df.mean(axis=1).fillna(0.0)
        benchmark_returns = self.close_df[self.benchmark].pct_change().fillna(
            0.0)

        return portfolio_returns, benchmark_returns
INFO - Baseline Metrics: Sharpe=-2.68, Return=-2.79%

Multi-Agent Persona Design and Debate Loop

When working through the Multi-Agent Persona Design and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

Engineering the Gemini Personas

When working through the Engineering the Gemini Personas stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

class Config:
    TICKERS = [
        "TSLA", "NVDA", "AMD", "PLTR", "COIN",
        "MARA", "RIOT", "MSTR", "HOOD", "ROKU",
        "UBER", "SHOP", "SNAP", "PTON", "DKNG"
    ]
    BENCHMARK = "SPY"
    MAX_ITERATIONS = 3
    INITIAL_PARAMS = {
        "lookback": 20,
        "atr_multiplier": 2.0,
        "volume_threshold": 1.5,
        "stop_loss_pct": 0.05
    }
class GeminiAgentPersona:
    def __init__(self, name: str, role: str, system_prompt: str):
        self.name = name
        self.role = role
        self.system_prompt = system_prompt
        self.client = genai.Client()

    def evaluate(self, current_params: dict, metrics: dict, iteration: int) -> dict:
        prompt = f"""
        You are {self.name}, a specialized {self.role} in a multi-agent quantitative trading system.
        Current Iteration: {iteration}
        Current Strategy Parameters: {json.dumps(current_params, indent=2)}
        Current Backtest Performance Metrics: {json.dumps(metrics, indent=2)}

        Analyze the performance and parameters from your professional perspective. Provide your critique and propose specific parameter adjustments (lookback, atr_multiplier, volume_threshold, stop_loss_pct) to improve the Sharpe ratio and reduce max drawdown.

        You MUST return a valid JSON object with keys:
        - "critique": Detailed analytical critique from your persona's perspective.
        - "proposed_adjustment": Dictionary of recommended parameter changes (e.g., ).
        - "confidence_score": Float between 0 and 1 representing your confidence in this adjustment.
        """

        response = self.client.models.generate_content(
            model="gemini-3.5-flash-lite",
            contents=prompt
        )

        text = response.text
        clean_text = text.strip()
        if clean_text.startswith("```json"):
            clean_text = clean_text[7:]
        if clean_text.endswith("```"):
            clean_text = clean_text[:-3]

        try:
            parsed = json.loads(clean_text.strip())
        except json.JSONDecodeError:
            parsed = {
                "critique": text[:500],
                "proposed_adjustment": current_params,
                "confidence_score": 0.6
            }
        return parsed

the author Iterative Debate Coordinator

When working through the the author Iterative Debate Coordinator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the the author Iterative Debate Coordinator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

class PersonaDebateCoordinator:
    def __init__(self, backtester: VolatilityStrategyBacktester, initial_params: dict, max_iterations: int):
        self.backtester = backtester
        self.current_params = initial_params.copy()
        self.max_iterations = max_iterations

        self.quant = GeminiAgentPersona(
            "QuantAnalyst",
            "Quantitative Strategist",
            "Focuses on alpha generation, indicator sensitivity, and statistical edge."
        )
        self.risk = GeminiAgentPersona(
            "RiskManager",
            "Risk Officer",
            "Focuses on drawdowns, tail risk, position sizing, and volatility caps."
        )
        self.execution = GeminiAgentPersona(
            "ExecutionStrategist",
            "Execution Trading Specialist",
            "Focuses on liquidity, volume confirmation, slippage, and practical execution friction."
        )

        self.audit_log = []
        self.parameter_history = []

    def run_iterative_loop(self) -> dict:
        best_sharpe = -999.0
        consensus_rationale = "Initial baseline strategy evaluation."

        p_ret, b_ret = self.backtester.backtest(self.current_params)
        baseline_metrics = calculate_performance_metrics(p_ret, b_ret)
        current_metrics = baseline_metrics

        logging.info(
            f"Baseline Metrics: Sharpe={baseline_metrics['sharpe_ratio']:.2f}, Return={baseline_metrics['cumulative_return']*100:.2f}%")

        for i in range(1, self.max_iterations + 1):
            logging.debug(
                f"--- Starting Debate Iteration {i}/{self.max_iterations} ---")

            q_eval = self.quant.evaluate(
                self.current_params, current_metrics, i)
            r_eval = self.risk.evaluate(
                self.current_params, current_metrics, i)
            e_eval = self.execution.evaluate(
                self.current_params, current_metrics, i)

            for persona_name, eval_data in [("QuantAnalyst", q_eval), ("RiskManager", r_eval), ("ExecutionStrategist", e_eval)]:
                self.audit_log.append({
                    "iteration_number": i,
                    "persona_name": persona_name,
                    "critique_text": eval_data.get("critique", ""),
                    "proposed_adjustment": eval_data.get("proposed_adjustment", {}),
                    "confidence_score": eval_data.get("confidence_score", 0.8)
                })

            q_adj = q_eval.get("proposed_adjustment", {})
            r_adj = r_eval.get("proposed_adjustment", {})
            e_adj = e_eval.get("proposed_adjustment", {})

            new_lookback = int(np.mean([
                float(q_adj.get("lookback",
                      self.current_params["lookback"])),
                float(r_adj.get("lookback",
                      self.current_params["lookback"])),
                float(e_adj.get("lookback",
                      self.current_params["lookback"]))
            ]))
            new_atr = float(np.mean([
                float(q_adj.get("atr_multiplier",
                      self.current_params["atr_multiplier"])),
                float(r_adj.get("atr_multiplier",
                      self.current_params["atr_multiplier"])),
                float(e_adj.get("atr_multiplier",
                      self.current_params["atr_multiplier"]))
            ]))
            new_vol_thresh = float(np.mean([
                float(q_adj.get("volume_threshold",
                      self.current_params["volume_threshold"])),
                float(r_adj.get("volume_threshold",
                      self.current_params["volume_threshold"])),
                float(e_adj.get("volume_threshold",
                      self.current_params["volume_threshold"]))
            ]))
            new_sl = float(np.mean([
                float(q_adj.get("stop_loss_pct",
                      self.current_params["stop_loss_pct"])),
                float(r_adj.get("stop_loss_pct",
                      self.current_params["stop_loss_pct"])),
                float(e_adj.get("stop_loss_pct",
                      self.current_params["stop_loss_pct"]))
            ]))

            self.current_params = {
                "lookback": max(5, min(60, new_lookback)),
                "atr_multiplier": max(1.0, min(4.0, new_atr)),
                "volume_threshold": max(1.0, min(3.0, new_vol_thresh)),
                "stop_loss_pct": max(0.01, min(0.15, new_sl))
            }

            p_ret_new, b_ret_new = self.backtester.backtest(
                self.current_params)
            current_metrics = calculate_performance_metrics(
                p_ret_new, b_ret_new)

            logging.debug(
                f"Iteration {i} Results: Sharpe={current_metrics['sharpe_ratio']:.2f}, Return={current_metrics['cumulative_return']*100:.2f}%, Params={self.current_params}")

            self.parameter_history.append({
                "iteration": i,
                "params": self.current_params.copy(),
                "metrics": current_metrics.copy()
            })

            if current_metrics["sharpe_ratio"] > best_sharpe:
                best_sharpe = current_metrics["sharpe_ratio"]
                consensus_rationale = f"Iteration {i} consensus achieved superior Sharpe ratio through optimized ATR multiplier and volume threshold refinement."

        summary_result = {
            "strategy_version": "v2.0-optimized",
            "final_sharpe_ratio": float(best_sharpe),
            "consensus_rationale": consensus_rationale,
            "max_drawdown_delta": float(current_metrics["max_drawdown"] - baseline_metrics["max_drawdown"]),
            "total_iterations_run": self.max_iterations,
            "parameter_history": self.parameter_history
        }

        return summary_result
INFO - Successfully generated output JSON files.

Results, Performance Analysis, and Benchmark Comparison

the author Results Performance Analysis and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Quantitative Breakdown of Strategy Iterations

the author Quantitative Breakdown of Strategy stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

def calculate_performance_metrics(portfolio_returns: pd.Series, benchmark_returns: pd.Series) -> dict:
    valid_idx = portfolio_returns.index.intersection(benchmark_returns.index)
    p_ret = portfolio_returns.loc[valid_idx]
    b_ret = benchmark_returns.loc[valid_idx]

    trading_days = 252
    cum_return = (1 + p_ret).prod() - 1
    annualized_return = (1 + cum_return) ** (trading_days /
                                             len(p_ret)) - 1 if len(p_ret) > 0 else 0.0
    annualized_vol = p_ret.std() * np.sqrt(trading_days)
    risk_free_rate = 0.03
    sharpe_ratio = (annualized_return - risk_free_rate) / \
        annualized_vol if annualized_vol > 0 else 0.0

    cum_wealth = (1 + p_ret).cumprod()
    peak = cum_wealth.cummax()
    drawdown = (cum_wealth - peak) / peak
    max_drawdown = drawdown.min()

    b_cum_return = (1 + b_ret).prod() - 1
    outperformance = cum_return - b_cum_return

    return {
        "cumulative_return": float(cum_return),
        "annualized_return": float(annualized_return),
        "annualized_volatility": float(annualized_vol),
        "sharpe_ratio": float(sharpe_ratio),
        "max_drawdown": float(max_drawdown),
        "benchmark_cumulative_return": float(b_cum_return),
        "outperformance": float(outperformance)
    }
{
    "strategy_version": "v2.0-optimized",
    "final_sharpe_ratio": -0.6822300207528283,
    "consensus_rationale": "Iteration 1 consensus achieved superior Sharpe ratio through optimized ATR multiplier and volume threshold refinement.",
    "parameter_history": [
        {
            "iteration": 1,
            "params": {
                "lookback": 14,
                "atr_multiplier": 1.5,
                "volume_threshold": 1.4666,
                "stop_loss_pct": 0.025
            },
            "metrics": {
                "cumulative_return": 0.0399,
                "annualized_return": 0.0132,
                "sharpe_ratio": -0.6822,
                "max_drawdown": -0.0344
            }
        }
    ]
}

Benchmark Realities and Strategy Visualization

the author Benchmark Realities and Strategy stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. the author Benchmark Realities and Strategy stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

def main():
    logging.info(
        "Starting Multi-Agent Volatility Breakout Optimization Pipeline.")

    loader = MarketDataLoader(Config.TICKERS, Config.BENCHMARK)
    close_df, high_df, low_df, volume_df = loader.fetch_data()

    backtester = VolatilityStrategyBacktester(
        close_df, high_df, low_df, volume_df, Config.BENCHMARK)

    # Run initial baseline backtest
    baseline_p_ret, baseline_b_ret = backtester.backtest(Config.INITIAL_PARAMS)

    coordinator = PersonaDebateCoordinator(
        backtester, Config.INITIAL_PARAMS, Config.MAX_ITERATIONS)
    summary_results = coordinator.run_iterative_loop()

    # Write output files exactly as specified
    with open("volatility_debate_summary.json", "w") as f:
        json.dump(summary_results, f, indent=4)

    with open("iteration_debate_audit_log.json", "w") as f:
        json.dump(coordinator.audit_log, f, indent=4)

    logging.info("Successfully generated output JSON files.")

    # Generate annotated plot
    fig, ax = plt.subplots(figsize=(10, 6))
    optimized_p_ret, _ = backtester.backtest(coordinator.current_params)

    wealth_baseline = (1 + baseline_p_ret).cumprod()
    wealth_optimized = (1 + optimized_p_ret).cumprod()
    wealth_benchmark = (1 + baseline_b_ret).cumprod()

    ax.plot(wealth_baseline.index, wealth_baseline,
            label="Baseline Strategy", color="gray", linestyle="--")
    ax.plot(wealth_optimized.index, wealth_optimized,
            label="Optimized Multi-Agent Strategy", color="blue", linewidth=2)
    ax.plot(wealth_benchmark.index, wealth_benchmark,
            label="SPY Benchmark", color="orange", alpha=0.7)

    # Annotate final peak / convergence
    max_idx = wealth_optimized.idxmax()
    max_val = wealth_optimized.max()
    ax.annotate(f"Optimized Peak: {max_val:.2f}x",
                xy=(max_idx, max_val),
                xytext=(max_idx - pd.Timedelta(days=90), max_val * 0.9),
                arrowprops=dict(facecolor='blue', shrink=0.05, width=1, headwidth=6))

    ax.set_title("Multi-Agent Volatility Breakout Portfolio vs. SPY Benchmark",
                 fontsize=12, fontweight='bold')
    ax.set_xlabel("Date", fontsize=10)
    ax.set_ylabel("Growth of 1.00", fontsize=10)
    ax.legend(loc="upper left")
    ax.grid(True, linestyle=":", alpha=0.6)

    plt.tight_layout()
    plt.show()
    logging.info("Successfully generated strategy performance plot.")
if __name__ == "__main__":
    main()
INFO - Successfully generated strategy performance plot.

Conclusion and Production Next Steps

For the Conclusion and Production Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 476dcdb15114: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.