Skip to content

GCP.Dataproc reference

Source: src/GCP/Dataproc/AutoscalingPolicy.ts

A Dataproc autoscaling policy (locations API).

Changing policyId or location replaces the policy. Worker bounds, algorithm, and labels update in place via a full-resource replace.

Generated name

const policy = yield* GCP.Dataproc.AutoscalingPolicy("SparkScale", {});

Explicit bounds

const policy = yield* GCP.Dataproc.AutoscalingPolicy("SparkScale", {
policyId: "spark-scale",
location: "us-central1",
workerConfig: { minInstances: 2, maxInstances: 6 },
labels: { env: "prod" },
});

Source: src/GCP/Dataproc/Batch.ts

A Dataproc serverless batch workload.

Batches are immutable after create. Changing identity or config replaces the batch (re-runs the workload). The resource exists as soon as the create operation returns; Spark completion is not required.

SparkPi

const batch = yield* GCP.Dataproc.Batch("Pi", {
sparkBatch: {
mainClass: "org.apache.spark.examples.SparkPi",
jarFileUris: ["file:///usr/lib/spark/examples/jars/spark-examples.jar"],
args: ["1"],
},
});

PySpark

const batch = yield* GCP.Dataproc.Batch("Etl", {
pysparkBatch: { mainPythonFileUri: "gs://bucket/job.py" },
labels: { env: "prod" },
});

Source: src/GCP/Dataproc/Cluster.ts

A Dataproc cluster of Compute Engine VMs.

Defaults to a single-node cluster (e2-standard-2, 30 GB boot disk) so a bare Cluster("Spark", {}) stays cheap. Set clusterType: "STANDARD" and workerNumInstances for a multi-VM cluster.

Changing clusterName, region, topology, image, machine types, disks, network, zone, or software config replaces the cluster. Labels, primary worker count, secondary worker count, and the autoscaling policy URI update in place.

Provisioning typically takes several minutes.

Generated name, single-node

const cluster = yield* GCP.Dataproc.Cluster("Spark", {});

Explicit id, labels, and machine type

const cluster = yield* GCP.Dataproc.Cluster("Spark", {
clusterName: "app-spark",
region: "us-central1",
clusterType: "SINGLE_NODE",
masterMachineType: "e2-standard-2",
masterBootDiskSizeGb: 30,
labels: { env: "prod" },
});
const cluster = yield* GCP.Dataproc.Cluster("Spark", {
clusterType: "STANDARD",
workerNumInstances: 2,
workerMachineType: "e2-standard-2",
workerBootDiskSizeGb: 30,
});

Source: src/GCP/Dataproc/GetCluster.ts

Runtime binding for Dataproc clusters.get.

Bind this operation to a Cluster in a Function/Action init phase. Provide GetClusterHttp.

const getCluster = yield* GCP.Dataproc.GetCluster(cluster);
const live = yield* getCluster();

Source: src/GCP/Dataproc/GetClusterHttp.ts Kind: Layer · Provides: GCP.Dataproc.GetCluster

HTTP implementation of GetCluster.

Source: src/GCP/Dataproc/RegionsAutoscalingPolicy.ts

A Dataproc autoscaling policy (regions API).

Same resource as AutoscalingPolicy addressed via projects/{project}/regions/{region}/autoscalingPolicies/{policy}.

RegionsAutoscalingPolicy: Creating a Policy

Section titled “RegionsAutoscalingPolicy: Creating a Policy”

Generated name

const policy = yield* GCP.Dataproc.RegionsAutoscalingPolicy("SparkScale", {});

Explicit bounds

const policy = yield* GCP.Dataproc.RegionsAutoscalingPolicy("SparkScale", {
policyId: "spark-scale-reg",
region: "us-central1",
workerConfig: { minInstances: 2, maxInstances: 6 },
labels: { env: "prod" },
});

Source: src/GCP/Dataproc/RegionsWorkflowTemplate.ts

A Dataproc workflow template (regions API).

Same resource as WorkflowTemplate addressed via projects/{project}/regions/{region}/workflowTemplates/{template}.

RegionsWorkflowTemplate: Creating a Template

Section titled “RegionsWorkflowTemplate: Creating a Template”

Generated name

const template = yield* GCP.Dataproc.RegionsWorkflowTemplate("Nightly", {});

Explicit id and labels

const template = yield* GCP.Dataproc.RegionsWorkflowTemplate("Nightly", {
templateId: "nightly-etl-reg",
region: "us-central1",
labels: { env: "prod" },
});

Source: src/GCP/Dataproc/Session.ts

A Dataproc serverless interactive session.

Sessions are immutable after create. Changing identity or config replaces the session. Provisioning typically takes several minutes.

Jupyter session

const session = yield* GCP.Dataproc.Session("Notebook", {
jupyterSession: { kernel: "PYTHON" },
environmentConfig: { executionConfig: { idleTtl: "600s", ttl: "3600s" } },
});

From a session template

const session = yield* GCP.Dataproc.Session("Notebook", {
sessionTemplate: template.name,
});

Source: src/GCP/Dataproc/SessionTemplate.ts

A Dataproc serverless interactive session template.

Changing templateId or location replaces the template. Description, labels, Jupyter config, runtime, and environment patch in place.

Generated name

const template = yield* GCP.Dataproc.SessionTemplate("Notebook", {});

Explicit Jupyter kernel

const template = yield* GCP.Dataproc.SessionTemplate("Notebook", {
templateId: "analytics-nb",
location: "us-central1",
description: "python notebooks",
jupyterSession: { kernel: "PYTHON" },
labels: { env: "prod" },
});

Source: src/GCP/Dataproc/SubmitJob.ts

Runtime binding for Dataproc jobs.submit.

Bind this operation to a Cluster in a Function/Action init phase. Provide SubmitJobHttp. The bound cluster is injected as job.placement.clusterName unless the request already sets one.

const submitJob = yield* GCP.Dataproc.SubmitJob(cluster);
const job = yield* submitJob({
body: {
job: {
sparkJob: {
mainClass: "org.apache.spark.examples.SparkPi",
jarFileUris: [
"file:///usr/lib/spark/examples/jars/spark-examples.jar",
],
args: ["10"],
},
},
},
});

Source: src/GCP/Dataproc/SubmitJobHttp.ts Kind: Layer · Provides: GCP.Dataproc.SubmitJob

HTTP implementation of SubmitJob.

Grants roles/dataproc.editor on the project because dataproc.jobs.create is checked on the project (jobs are not children of the cluster, so neither a cluster-level grant nor an IAM Condition on the cluster name applies) and no narrower predefined role contains it.

Source: src/GCP/Dataproc/WorkflowTemplate.ts

A Dataproc workflow template (locations API).

Stores a DAG of jobs and a cluster placement. Instantiating the template (out of band) is what actually creates a cluster and runs jobs.

Changing templateId or location replaces the template. Jobs, placement, labels, and timeout update in place (optimistic version).

Generated name

const template = yield* GCP.Dataproc.WorkflowTemplate("Nightly", {});

Explicit jobs

const template = yield* GCP.Dataproc.WorkflowTemplate("Nightly", {
templateId: "nightly-etl",
location: "us-central1",
labels: { env: "prod" },
jobs: [{
stepId: "spark-pi",
sparkJob: {
mainClass: "org.apache.spark.examples.SparkPi",
jarFileUris: ["file:///usr/lib/spark/examples/jars/spark-examples.jar"],
},
}],
});