Home / Articles / Practical notes: Debugging Serverless Apache Spark using Gemini with MCP

This article is published in English.

Practical notes: Debugging Serverless Apache Spark using Gemini with MCP

Operable walkthrough of Practical notes: Debugging Serverless Apache Spark using Gemini with MCP: contracts, checks, and drop-in code slots for teams shipping this pattern.

1385 words

This walkthrough rebuilds the path from raw materials to a working system for: Debugging Serverless Apache Spark using Gemini with MCP. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

The limits of zero-context AI

When working through the The limits of zero-context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

py4j.protocol.Py4JJavaError: An error occurred while calling o80.load.
org.apache.spark.SparkException: Job aborted due to stage failure.
Traceback (most recent call last):
  File "spark_job.py", line 26, in main
    df_with_status = df.withColunm("status", lit("active"))
AttributeError: 'DataFrame' object has no attribute 'withColunm'

Bringing context to your terminal with Google Antigravity CLI

When working through the Bringing context to your stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

export GOOGLE_CLOUD_PROJECT="$PROJECT_ID"
mkdir -p ~/.gemini/antigravity-cli
cat << 'EOF' | tee ~/.gemini/antigravity-cli/settings.json ~/.gemini/jetski/cli/settings.json ~/.gemini/antigravity/settings.json ~/.gemini/settings.json
{
  "toolPermission": "always-proceed",
  "permissions": {
    "allow": [
      "read_file(*)",
      "write_file(*)",
      "mcp(*)"
    ]
  }
}
EOF
agy -p "examine spark_job.py, fix the DataFrame method typo, and save the file"

Full-infrastructure debugging with the Spark MCP server

When working through the Full-infrastructure debugging with the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the Full-infrastructure debugging with the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

mkdir -p ~/.gemini/config
cat << EOF | tee ~/.gemini/config/mcp_config.json
{
  "mcpServers": {
    "spark": {
      "serverUrl": "https://dataproc-${REGION}.googleapis.com/mcp"
    }
  }
}
EOF
agy -p "inspect my latest failed Spark batch via MCP, identify the root cause, and fix spark_job.py so it succeeds"

Codifying playbooks with Agent Skills

The Codifying playbooks with Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

mkdir -p .agents/skills/spark-troubleshooter
cat << 'EOF' > .agents/skills/spark-troubleshooter/SKILL.md
---
name: spark-troubleshooter
description: Diagnoses failed Apache Spark batches on Managed Service for Apache Spark, inspects live batch logs via the Spark MCP server, and recommends resolution commands. Use when troubleshooting Spark job failures.
---

# Spark Troubleshooter Skill

This skill diagnoses failed Apache Spark batches on Managed Service for Apache Spark and offers rapid solutions.

## Instructions
1. Verify local syntax: Locate the PySpark script in the current directory and check for compilation or syntax issues.
2. Fetch live batch state: Call the Spark MCP server tool list_batches and inspect the status of the most recent batch.
3. Check JVM and PySpark logs: Look for standard Spark exceptions, such as FileNotFoundException, AnalysisException, or out-of-memory errors in the batch logs.
4. Recommend action:
    * If a Cloud Storage path is invalid, recommend the exact gcloud storage buckets create command or update the script path.
    * If the job fails due to configuration, generate the correct gcloud dataproc batches submit command with the appropriate parameters.
    * If a syntax error is detected, fix the code in-place.
    * For other errors, recommend a fix.
EOF
agy -p "Diagnose why my last Spark batch failed and recommend a fix"
[Spark Troubleshooter] Running diagnostic playbook...
- Local Syntax: OK (spark_job.py has valid python syntax)
- Spark Batch Status: FAILED (batch-928f1)
- Log Exception: java.io.FileNotFoundException for gs://my-missing-bucket/input.csv

Recommendation:
The bucket gs://my-missing-bucket does not exist. Run this command to create it:

  gcloud storage buckets create gs://my-missing-bucket --location=$REGION

Once created, submit the batch again with:

  gcloud dataproc batches submit pyspark spark_job.py \
      --region=$REGION \
      --deps-bucket=gs://$BUCKET_NAME

Web-based troubleshooting with Gemini in Cloud Logging

The Web-based troubleshooting with Gemini stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Summary

The Summary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The Summary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 3af041a23886: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.