GCP.Dataproc reference
AutoscalingPolicy
Section titled “AutoscalingPolicy”Source:
src/GCP/Dataproc/AutoscalingPolicy.ts
A Dataproc autoscaling policy (locations API).
Changing policyId or location replaces the policy. Worker bounds,
algorithm, and labels update in place via a full-resource replace.
AutoscalingPolicy: Creating a Policy
Section titled “AutoscalingPolicy: Creating a Policy”Generated name
const policy = yield* GCP.Dataproc.AutoscalingPolicy("SparkScale", {});Explicit bounds
const policy = yield* GCP.Dataproc.AutoscalingPolicy("SparkScale", { policyId: "spark-scale", location: "us-central1", workerConfig: { minInstances: 2, maxInstances: 6 }, labels: { env: "prod" },});Source:
src/GCP/Dataproc/Batch.ts
A Dataproc serverless batch workload.
Batches are immutable after create. Changing identity or config replaces the batch (re-runs the workload). The resource exists as soon as the create operation returns; Spark completion is not required.
Batch: Creating a Batch
Section titled “Batch: Creating a Batch”SparkPi
const batch = yield* GCP.Dataproc.Batch("Pi", { sparkBatch: { mainClass: "org.apache.spark.examples.SparkPi", jarFileUris: ["file:///usr/lib/spark/examples/jars/spark-examples.jar"], args: ["1"], },});PySpark
const batch = yield* GCP.Dataproc.Batch("Etl", { pysparkBatch: { mainPythonFileUri: "gs://bucket/job.py" }, labels: { env: "prod" },});Cluster
Section titled “Cluster”Source:
src/GCP/Dataproc/Cluster.ts
A Dataproc cluster of Compute Engine VMs.
Defaults to a single-node cluster (e2-standard-2, 30 GB boot disk) so
a bare Cluster("Spark", {}) stays cheap. Set clusterType: "STANDARD"
and workerNumInstances for a multi-VM cluster.
Changing clusterName, region, topology, image, machine types, disks,
network, zone, or software config replaces the cluster. Labels, primary
worker count, secondary worker count, and the autoscaling policy URI
update in place.
Provisioning typically takes several minutes.
Cluster: Creating a Cluster
Section titled “Cluster: Creating a Cluster”Generated name, single-node
const cluster = yield* GCP.Dataproc.Cluster("Spark", {});Explicit id, labels, and machine type
const cluster = yield* GCP.Dataproc.Cluster("Spark", { clusterName: "app-spark", region: "us-central1", clusterType: "SINGLE_NODE", masterMachineType: "e2-standard-2", masterBootDiskSizeGb: 30, labels: { env: "prod" },});Cluster: Standard topology
Section titled “Cluster: Standard topology”const cluster = yield* GCP.Dataproc.Cluster("Spark", { clusterType: "STANDARD", workerNumInstances: 2, workerMachineType: "e2-standard-2", workerBootDiskSizeGb: 30,});GetCluster
Section titled “GetCluster”Source:
src/GCP/Dataproc/GetCluster.ts
Runtime binding for Dataproc clusters.get.
Bind this operation to a Cluster in a Function/Action init
phase. Provide GetClusterHttp.
GetCluster: Observing Clusters
Section titled “GetCluster: Observing Clusters”const getCluster = yield* GCP.Dataproc.GetCluster(cluster);const live = yield* getCluster();GetClusterHttp
Section titled “GetClusterHttp”Source:
src/GCP/Dataproc/GetClusterHttp.tsKind: Layer · Provides:GCP.Dataproc.GetCluster
HTTP implementation of GetCluster.
RegionsAutoscalingPolicy
Section titled “RegionsAutoscalingPolicy”Source:
src/GCP/Dataproc/RegionsAutoscalingPolicy.ts
A Dataproc autoscaling policy (regions API).
Same resource as AutoscalingPolicy addressed via
projects/{project}/regions/{region}/autoscalingPolicies/{policy}.
RegionsAutoscalingPolicy: Creating a Policy
Section titled “RegionsAutoscalingPolicy: Creating a Policy”Generated name
const policy = yield* GCP.Dataproc.RegionsAutoscalingPolicy("SparkScale", {});Explicit bounds
const policy = yield* GCP.Dataproc.RegionsAutoscalingPolicy("SparkScale", { policyId: "spark-scale-reg", region: "us-central1", workerConfig: { minInstances: 2, maxInstances: 6 }, labels: { env: "prod" },});RegionsWorkflowTemplate
Section titled “RegionsWorkflowTemplate”Source:
src/GCP/Dataproc/RegionsWorkflowTemplate.ts
A Dataproc workflow template (regions API).
Same resource as WorkflowTemplate addressed via
projects/{project}/regions/{region}/workflowTemplates/{template}.
RegionsWorkflowTemplate: Creating a Template
Section titled “RegionsWorkflowTemplate: Creating a Template”Generated name
const template = yield* GCP.Dataproc.RegionsWorkflowTemplate("Nightly", {});Explicit id and labels
const template = yield* GCP.Dataproc.RegionsWorkflowTemplate("Nightly", { templateId: "nightly-etl-reg", region: "us-central1", labels: { env: "prod" },});Session
Section titled “Session”Source:
src/GCP/Dataproc/Session.ts
A Dataproc serverless interactive session.
Sessions are immutable after create. Changing identity or config replaces the session. Provisioning typically takes several minutes.
Session: Creating a Session
Section titled “Session: Creating a Session”Jupyter session
const session = yield* GCP.Dataproc.Session("Notebook", { jupyterSession: { kernel: "PYTHON" }, environmentConfig: { executionConfig: { idleTtl: "600s", ttl: "3600s" } },});From a session template
const session = yield* GCP.Dataproc.Session("Notebook", { sessionTemplate: template.name,});SessionTemplate
Section titled “SessionTemplate”Source:
src/GCP/Dataproc/SessionTemplate.ts
A Dataproc serverless interactive session template.
Changing templateId or location replaces the template. Description,
labels, Jupyter config, runtime, and environment patch in place.
SessionTemplate: Creating a Template
Section titled “SessionTemplate: Creating a Template”Generated name
const template = yield* GCP.Dataproc.SessionTemplate("Notebook", {});Explicit Jupyter kernel
const template = yield* GCP.Dataproc.SessionTemplate("Notebook", { templateId: "analytics-nb", location: "us-central1", description: "python notebooks", jupyterSession: { kernel: "PYTHON" }, labels: { env: "prod" },});SubmitJob
Section titled “SubmitJob”Source:
src/GCP/Dataproc/SubmitJob.ts
Runtime binding for Dataproc jobs.submit.
Bind this operation to a Cluster in a Function/Action init
phase. Provide SubmitJobHttp. The bound cluster is injected as
job.placement.clusterName unless the request already sets one.
SubmitJob: Submitting Jobs
Section titled “SubmitJob: Submitting Jobs”const submitJob = yield* GCP.Dataproc.SubmitJob(cluster);const job = yield* submitJob({ body: { job: { sparkJob: { mainClass: "org.apache.spark.examples.SparkPi", jarFileUris: [ "file:///usr/lib/spark/examples/jars/spark-examples.jar", ], args: ["10"], }, }, },});SubmitJobHttp
Section titled “SubmitJobHttp”Source:
src/GCP/Dataproc/SubmitJobHttp.tsKind: Layer · Provides:GCP.Dataproc.SubmitJob
HTTP implementation of SubmitJob.
Grants roles/dataproc.editor on the project because
dataproc.jobs.create is checked on the project (jobs are not children
of the cluster, so neither a cluster-level grant nor an IAM Condition on
the cluster name applies) and no narrower predefined role contains it.
WorkflowTemplate
Section titled “WorkflowTemplate”Source:
src/GCP/Dataproc/WorkflowTemplate.ts
A Dataproc workflow template (locations API).
Stores a DAG of jobs and a cluster placement. Instantiating the template (out of band) is what actually creates a cluster and runs jobs.
Changing templateId or location replaces the template. Jobs,
placement, labels, and timeout update in place (optimistic version).
WorkflowTemplate: Creating a Template
Section titled “WorkflowTemplate: Creating a Template”Generated name
const template = yield* GCP.Dataproc.WorkflowTemplate("Nightly", {});Explicit jobs
const template = yield* GCP.Dataproc.WorkflowTemplate("Nightly", { templateId: "nightly-etl", location: "us-central1", labels: { env: "prod" }, jobs: [{ stepId: "spark-pi", sparkJob: { mainClass: "org.apache.spark.examples.SparkPi", jarFileUris: ["file:///usr/lib/spark/examples/jars/spark-examples.jar"], }, }],});