Professional-Data-Engineer Questions and Answers

Question # 6

If you want to create a machine learning model that predicts the price of a particular stock based on its recent price history, what type of estimator should you use?

Unsupervised learning

Regressor

Classifier

Clustering estimator

Full Access

Question # 7

You use BigQuery as your centralized analytics platform. New data is loaded every day, and an ETL pipeline modifies the original data and prepares it for the final users. This ETL pipeline is regularly modified and can generate errors, but sometimes the errors are detected only after 2 weeks. You need to provide a method to recover from these errors, and your backups should be optimized for storage costs. How should you organize your data in BigQuery and store your backups?

Organize your data in a single table, export, and compress and store the BigQuery data in Cloud Storage.

Organize your data in separate tables for each month, and export, compress, and store the data in Cloud Storage.

Organize your data in separate tables for each month, and duplicate your data on a separate dataset in BigQuery.

Organize your data in separate tables for each month, and use snapshot decorators to restore the table to a time prior to the corruption.

Full Access

Question # 8

Flowlogistic wants to use Google BigQuery as their primary analysis system, but they still have Apache Hadoop and Spark workloads that they cannot move to BigQuery. Flowlogistic does not know how to store the data that is common to both workloads. What should they do?

Store the common data in BigQuery as partitioned tables.

Store the common data in BigQuery and expose authorized views.

Store the common data encoded as Avro in Google Cloud Storage.

Store he common data in the HDFS storage for a Google Cloud Dataproc cluster.

Full Access

Question # 9

Flowlogistic’s CEO wants to gain rapid insight into their customer base so his sales team can be better informed in the field. This team is not very technical, so they’ve purchased a visualization tool to simplify the creation of BigQuery reports. However, they’ve been overwhelmed by all thedata in the table, and are spending a lot of money on queries trying to find the data they need. You want to solve their problem in the most cost-effective way. What should you do?

Export the data into a Google Sheet for virtualization.

Create an additional table with only the necessary columns.

Create a view on the table to present to the virtualization tool.

Create identity and access management (IAM) roles on the appropriate columns, so only they appear in a query.

Full Access

Question # 10

Flowlogistic is rolling out their real-time inventory tracking system. The tracking devices will all send package-tracking messages, which will now go to a single Google Cloud Pub/Sub topic instead of the Apache Kafka cluster. A subscriber application will then process the messages for real-time reporting and store them in Google BigQuery for historical analysis. You want to ensure the package data can be analyzed over time.

Which approach should you take?

Attach the timestamp on each message in the Cloud Pub/Sub subscriber application as they are received.

Attach the timestamp and Package ID on the outbound message from each publisher device as they are sent to Clod Pub/Sub.

Use the NOW () function in BigQuery to record the event’s time.

Use the automatically generated timestamp from Cloud Pub/Sub to order the data.

Full Access

Question # 11

Which of these operations can you perform from the BigQuery Web UI?

Upload a file in SQL format.

Load data with nested and repeated fields.

Upload a 20 MB file.

Upload multiple files using a wildcard.

Full Access

Question # 12

Flowlogistic’s management has determined that the current Apache Kafka servers cannot handle the data volume for their real-time inventory tracking system. You need to build a new system on Google Cloud Platform (GCP) that will feed the proprietary tracking software. The system must be able to ingest data from a variety of global sources, process and query in real-time, and store the data reliably. Which combination of GCP products should you choose?

Cloud Pub/Sub, Cloud Dataflow, and Cloud Storage

Cloud Pub/Sub, Cloud Dataflow, and Local SSD

Cloud Pub/Sub, Cloud SQL, and Cloud Storage

Cloud Load Balancing, Cloud Dataflow, and Cloud Storage

Full Access

Question # 13

You are designing a system that requires an ACID-compliant database. You must ensure that the system requires minimal human intervention in case of a failure. What should you do?

Configure a Cloud SQL for MySQL instance with point-in-time recovery enabled.

Configure a Cloud SQL for PostgreSQL instance with high availability enabled.

Configure a Bigtable instance with more than one cluster.

Configure a BJgQuery table with a multi-region configuration.

Full Access

Question # 14

You are developing a model to identify the factors that lead to sales conversions for your customers. You have completed processing your data. You want to continue through the model development lifecycle. What should you do next?

Use your model to run predictions on fresh customer input data.

Test and evaluate your model on your curated data to determine how well the model performs.

Monitor your model performance, and make any adjustments needed.

Delineate what data will be used for testing and what will be used for training the model.

Full Access

Question # 15

You are collecting loT sensor data from millions of devices across the world and storing the data in BigQuery. Your access pattern is based on recent data tittered by location_id and device_version with the following query:

You want to optimize your queries for cost and performance. How should you structure your data?

Partition table data by create_date, location_id and device_version

Partition table data by create_date cluster table data by tocation_id and device_version

Cluster table data by create_date location_id and device_version

Cluster table data by create_date, partition by location and device_version

Full Access

Question # 16

You used Cloud Dataprep to create a recipe on a sample of data in a BigQuery table. You want to reuse this recipe on a daily upload of data with the same schema, after the load job with variable execution time completes. What should you do?

Create a cron schedule in Cloud Dataprep.

Create an App Engine cron job to schedule the execution of the Cloud Dataprep job.

Export the recipe as a Cloud Dataprep template, and create a job in Cloud Scheduler.

Export the Cloud Dataprep job as a Cloud Dataflow template, and incorporate it into a Cloud Composer job.

Full Access

Question # 17

You maintain ETL pipelines. You notice that a streaming pipeline running on Dataflow is taking a long time to process incoming data, which causes output delays. You also noticed that the pipeline graph was automatically optimized by Dataflow and merged into one step. You want to identify where the potential bottleneck is occurring. What should you do?

Insert a Reshuffle operation after each processing step, and monitor the execution details in the Dataflow console.

Log debug information in each ParDo function, and analyze the logs at execution time.

Insert output sinks after each key processing step, and observe the writing throughput of each block.

Verify that the Dataflow service accounts have appropriate permissions to write the processed data to the output sinks

Full Access

Answer:

Explanation:

When Dataflow fuses multiple transformations into a single stage (step), it can make it harder to pinpoint which specific part of that fused stage is causing a bottleneck because internal metrics for individual ParDos within the fused stage might not be as distinct.

Reshuffle Operation (Option D):Inserting a Reshuffle (or GroupByKey followed by ungrouping, which forces a shuffle) operation between logical processing steps in your Beam pipeline prevents Dataflow from fusing those steps. A shuffle operation acts as a barrier to fusion. This materializes the intermediate PCollection and forces data to be redistributed across workers.

Benefit for Debugging:By breaking the fusion, the Dataflow monitoring UI will display distinct steps for the operations before and after the Reshuffle. This allows you to observe metrics like processing time, throughput, and watermarks for each now-separated step, making it much easier to identify which part of your original fused logic is the bottleneck.

Let's analyze why other options are less effective for this specific problem of afused step:

A (Verify service account permissions):While important for overall pipeline health, permission issues usually result in outright failures or errors in logs, not typically a slowdown within a successfully running (albeit slow) fused step.

B (Insert output sinks):Adding actual output sinks (like writing to Pub/Sub or GCS) after each key step would also break fusion and allow you to measure throughput. However, it's a more heavyweight approach than Reshuffle. It introduces I/O overhead and requires setting up and managing these temporary sinks. Reshuffle is a lighter-weight way to achieve the same goal of breaking fusion for diagnostic purposes within the pipeline itself.

C (Log debug information):Logging can be helpful, but if the entire fused step is slow, logs might not easily distinguish which internal operation is the culprit without very careful and verbose logging. Analyzing potentially massive volumes of logs for performance bottlenecks can be less direct than observing stage metrics in the Dataflow UI once fusion is broken.

Using Reshuffle is a standard technique recommended by Google Cloud for debugging performance issues in fused Dataflow stages.

[Reference:, Google Cloud Documentation: Dataflow > Troubleshooting Dataflow pipelines > Common Dataflow errors and troubleshooting steps > Pipeline is slow or stuck. "Break transform fusion: Certain transforms in your pipeline might be fused together into a single stage for optimization. If a particular fused stage is causing a bottleneck, you can temporarily add Reshuffle transforms between the fused transforms to break them into smaller, separate stages. This allows you to get more visibility into the performance of each individual transform and isolate the bottleneck.", Apache Beam Documentation: Programming Guide > Pipeline I/O > Reshuffle."Reshuffle can be used to prevent fusion, and ensure that data is materialized and redistributed." (While the primary purpose of Reshuffle is often related to data distribution and freshness, a side effect and common use case is to break fusion for monitoring and debugging)., , , , ]

Question # 18

You are designing a messaging system by using Pub/Sub to process clickstream data with an event-driven consumer app that relies on a push subscription. You need to configure the messaging system that is reliable enough to handle temporary downtime of the consumer app. You also need the messaging system to store the input messages that cannot be consumed by the subscriber. The system needs to retry failed messages gradually, avoiding overloading the consumer app, and store the failed messages after a maximum of 10 retries in a topic. How should you configure the Pub/Sub subscription?

Increase the acknowledgement deadline to 10 minutes.

Use immediate redelivery as the subscription retry policy, and configure dead lettering to a different topic with maximum delivery attempts set to 10.

Use exponential backoff as the subscription retry policy, and configure dead lettering to the same source topic with maximum delivery attempts set to 10.

Use exponential backoff as the subscription retry policy, and configure dead lettering to a different topic with maximum delivery attempts set to 10.

Full Access

Question # 19

What are two methods that can be used to denormalize tables in BigQuery?

1) Split table into multiple tables; 2) Use a partitioned table

1) Join tables into one table; 2) Use nested repeated fields

1) Use a partitioned table; 2) Join tables into one table

1) Use nested repeated fields; 2) Use a partitioned table

Full Access

Question # 20

You are managing a Cloud Dataproc cluster. You need to make a job run faster while minimizing costs, without losing work in progress on your clusters. What should you do?

Increase the cluster size with more non-preemptible workers.

Increase the cluster size with preemptible worker nodes, and configure them to forcefully decommission.

Increase the cluster size with preemptible worker nodes, and use Cloud Stackdriver to trigger a script to preserve work.

Increase the cluster size with preemptible worker nodes, and configure them to use graceful decommissioning.

Full Access

Question # 21

Which of the following statements about Legacy SQL and Standard SQL is not true?

Standard SQL is the preferred query language for BigQuery.

If you write a query in Legacy SQL, it might generate an error if you try to run it with Standard SQL.

One difference between the two query languages is how you specify fully-qualified table names (i.e. table names that include their associated project name).

You need to set a query language for each dataset and the default is Standard SQL.

Full Access

Question # 22

Which of the following statements about the Wide & Deep Learning model are true? (Select 2 answers.)

The wide model is used for memorization, while the deep model is used for generalization.

A good use for the wide and deep model is a recommender system.

The wide model is used for generalization, while the deep model is used for memorization.

A good use for the wide and deep model is a small-scale linear regression problem.

Full Access

Question # 23

Which of the following statements is NOT true regarding Bigtable access roles?

Using IAM roles, you cannot give a user access to only one table in a project, rather than all tables in a project.

To give a user access to only one table in a project, grant the user the Bigtable Editor role forthat table.

You can configure access control only at the project level.

To give a user access to only one table in a project, you must configure access through your application.

Full Access

Question # 24

Government regulations in the banking industry mandate the protection of client’s personally identifiable information (PII). Your company requires PII to be access controlled encrypted and compliant with major data protection standards In addition to using Cloud Data Loss Prevention (Cloud DIP) you want to follow Google-recommended practices and use service accounts to control access to PII. What should you do?

Assign the required identity and Access Management (IAM) roles to every employee, and create a single service account to access protect resources

Use one service account to access a Cloud SQL database and use separate service accounts for each human user

Use Cloud Storage to comply with major data protection standards. Use one service account shared by all users

Use Cloud Storage to comply with major data protection standards. Use multiple service accounts attached to IAM groups to grant the appropriate access to each group

Full Access

Question # 25

Your company receives both batch- and stream-based event data. You want to process the data using Google Cloud Dataflow over a predictable time period. However, you realize that in some instances data can arrive late or out of order. How should you design your Cloud Dataflow pipeline to handle data that is late or out of order?

Set a single global window to capture all the data.

Set sliding windows to capture all the lagged data.

Use watermarks and timestamps to capture the lagged data.

Ensure every datasource type (stream or batch) has a timestamp, and use the timestamps to define the logic for lagged data.

Full Access

Question # 26

Your startup has a web application that currently serves customers out of a single region in Asia. You are targeting funding that will allow your startup lo serve customers globally. Your current goal is to optimize for cost, and your post-funding goat is to optimize for global presence and performance. You must use a native JDBC driver. What should you do?

Use Cloud Spanner to configure a single region instance initially. and then configure multi-region C oud Spanner instances after securing funding.

Use a Cloud SQL for PostgreSQL highly available instance first, and 8»gtable with US. Europe, and Asiareplication alter securing funding

Use a Cloud SQL for PostgreSQL zonal instance first and Bigtable with US. Europe, and Asia after securing funding.

Use a Cloud SOL for PostgreSQL zonal instance first, and Cloud SOL for PostgreSQL with highly available configuration after securing funding.

Full Access

Question # 27

You have a data pipeline with a Cloud Dataflow job that aggregates and writes time series metrics to Cloud Bigtable. This data feeds a dashboard used by thousands of users across the organization. You need to support additional concurrent users and reduce the amount of time required to write the data. Which two actions should you take? (Choose two.)

Configure your Cloud Dataflow pipeline to use local execution

Increase the maximum number of Cloud Dataflow workers by setting maxNumWorkers in PipelineOptions

Increase the number of nodes in the Cloud Bigtable cluster

Modify your Cloud Dataflow pipeline to use the Flatten transform before writing to Cloud Bigtable

Modify your Cloud Dataflow pipeline to use the CoGroupByKey transform before writing to Cloud Bigtable

Full Access

Question # 28

Your team has created several BigQuery curated datasets containing anonymized industry benchmark data. You want to make these datasets easily discoverable and accessible for querying by external partner companies within their own Google Cloud projects. You need a secure and scalable solution. What should you do?

Grant the roles/bigquery.dataViewer IAM role to the partner group email addresses on the datasets.

Publish the datasets as listings within BigQuery sharing (Analytics Hub).

Create authorized views for each dataset and grant access to each partner.

Export the datasets to partner-specific Cloud Storage buckets.

Full Access

Question # 29

Which of the following is not possible using primitive roles?

Give a user viewer access to BigQuery and owner access to Google Compute Engine instances.

Give UserA owner access and UserB editor access for all datasets in a project.

Give a user access to view all datasets in a project, but not run queries on them.

Give GroupA owner access and GroupB editor access for all datasets in a project.

Full Access

Question # 30

Which of these statements about BigQuery caching is true?

By default, a query's results are not cached.

BigQuery caches query results for 48 hours.

Query results are cached even if you specify a destination table.

There is no charge for a query that retrieves its results from cache.

Full Access

Question # 31

What are the minimum permissions needed for a service account used with Google Dataproc?

Execute to Google Cloud Storage; write to Google Cloud Logging

Write to Google Cloud Storage; read to Google Cloud Logging

Execute to Google Cloud Storage; execute to Google Cloud Logging

Read and write to Google Cloud Storage; write to Google Cloud Logging

Full Access

Question # 32

To run a TensorFlow training job on your own computer using Cloud Machine Learning Engine, what would your command start with?

gcloud ml-engine local train

gcloud ml-engine jobs submit training

gcloud ml-engine jobs submit training local

You can't run a TensorFlow program on your own computer using Cloud ML Engine .

Full Access

Question # 33

Which of the following is not true about Dataflow pipelines?

Pipelines are a set of operations

Pipelines represent a data processing job

Pipelines represent a directed graph of steps

Pipelines can share data between instances

Full Access

Question # 34

Your company has a hybrid cloud initiative. You have a complex data pipeline that moves data between cloud provider services and leverages services from each of the cloud providers. Which cloud-native service should you use to orchestrate the entire pipeline?

Cloud Dataflow

Cloud Composer

Cloud Dataprep

Cloud Dataproc

Full Access

Question # 35

You manage your company's BigQuery data warehouse. You need to implement a solution that enables the data science team to modify data for experiments without affecting the original tables, while minimizing additional storage costs. What should you do?

Set up authorized views in a shared dataset that reference the original tables.

Create snapshots of all the tables and restore them for the data science team to use.

Create table clones of all the tables for the data science team to use.

Create a separate dataset with full copies of all the tables for each member of the data science team.

Full Access

Question # 36

You are planning to use Cloud Storage as pad of your data lake solution. The Cloud Storage bucket will contain objects ingested from external systems. Each object will be ingested once, and the access patterns of individual objects will be random. You want to minimize the cost of storing and retrieving these objects. You want to ensure that any cost optimization efforts are transparent to the users and applications. What should you do?

Create a Cloud Storage bucket with Autoclass enabled.

Create a Cloud Storage bucket with an Object Lifecycle Management policy to transition objects from Standard to Coldline storage class if an object age reaches 30 days.

Create a Cloud Storage bucket with an Object Lifecycle Management policy to transition objects from Standard to Coldline storage class if an object is not live.

Create two Cloud Storage buckets. Use the Standard storage class for the first bucket, and use the Coldline storage class for the second bucket. Migrate objects from the first bucket to the second bucket after 30 days.

Full Access

Question # 37

Your company has data assets across multiple Cloud Storage buckets and BigQuery datasets containing raw and processed data. The requirement is to establish a unified data governance framework that allows for centralized metadata discovery, data quality monitoring, and consistent security policy application across these various data stores without physically moving or duplicating the data. You need to implement a solution to achieve this federated governance. What should you do?

Deploy a centralized Cloud SQL database to store metadata extracted from BigQuery and Cloud Storage using custom scripts.

Integrate the database with Looker Studio for data discovery and visualization.

Implement a custom policy engine using Cloud Run functions triggered by changes in IAM policies to enforce consistent security across projects.

Create a Looker Studio dashboard on BigQuery INFORMATION_SCHEMA views to visualize and monitor data quality.

Manage security using IAM policies at the project level, supplemented by BigQuery authorized views for granular access control.

Export metadata out of Dataplex Universal Catalog by running a metadata export job.

Implement Dataproc Metastore to manage table schemas and Apache Hive metastore for metadata discovery.

Manage security using a combination of BigQuery row-level security and Cloud Storage policies.

Use Dataplex to organize the BigQuery datasets and Cloud Storage buckets into lakes and zones.

Use Dataplex for automated metadata discovery, centralized security policy management, data profiling, and data quality tasks.

Full Access

Answer:

Explanation:

Dataplex is Google Cloud's intelligent data fabric that enables organizations to centrally manage, monitor, and govern data across data lakes, data warehouses, and data marts. It is specifically designed to provide a unified interface for distributed data without moving the data.

Unified Governance: Dataplex allows you to group distributed data (in GCS and BigQuery) into logical Lakes and Zones (e.g., Raw vs. Curated). This structure enables you to apply security policies once at the Lake or Zone level, which then propagates to all underlying assets.

Metadata Discovery: Dataplex automatically scans and registers metadata from your buckets and datasets into the Search/Catalog interface, making it discoverable without manual scripts.

Data Quality & Profiling: Dataplex includes built-in, fully managed features for Data Quality (declarative rules) and Data Profiling to monitor the health of your data directly within the governance framework.

Correcting other options:

A & B: These involve significant manual overhead, custom scripts, and fragmented tools. They do not provide a "unified data governance framework" that scales naturally across both GCS and BigQuery.

C: Dataproc Metastore is primarily for Hive/Spark workloads and doesn't offer the comprehensive security and data quality governance required for a general-purpose federated framework across BigQuery and Cloud Storage.

[Reference: Google Cloud Documentation on Dataplex:, "Dataplex is an intelligent data fabric that helps you unify distributed data and automate data management and governance across that data. With Dataplex, you can: Build a unified search and discovery experience across all your data. Centrally manage security and governance across data stored in Cloud Storage and BigQuery. Automate data quality and data profiling to ensure the reliability of your data assets." (Source: Dataplex Overview), "Dataplex lets you organize your data into lakes and zones. A lake represents a logical data domain... A zone represents a sub-domain within a lake and is useful for categorizing data by its readiness (e.g., raw vs. curated)." (Source: Dataplex terminology), , , ]

Question # 38

Your financial services company has a critical daily reconciliation process that involves several distinct steps: fetching data from an external SFTP server, decrypting the files, loading them into Cloud Storage, and finally running a series of BigQuery SQL transformations. Each step has strict dependencies, and the entire process should notify you if not completed by 7:00 AM. Manual intervention for failures is costly and delays compliance reporting. You need a highly observable and robust solution that supports easy re-runs of individual steps if errors occur. What should you do?

Define a Cloud Composer DAG to orchestrate the SFTP fetch and decryption steps, and then use Cloud Scheduler to trigger a separate Dataflow job that handles the Cloud Storage load and BigQuery transformations and schedule to run daily.

Develop a Cloud Composer DAG that includes a single PythonOperator to execute a Python script that runs each step sequentially, incorporating error handling and retries. Upload the scripts to a Cloud Composer environment's DAGs folder, and configure it to run daily.

Create a Cloud Composer DAG that includes a single BashOperator to execute a top-level shell script, which in turn calls individual scripts for each pipeline step. Upload the scripts to a Cloud Composer environment's DAGs folder, and configure it to run daily.

Implement a Cloud Composer DAG, with each step defined as a separate task using appropriate Airflow operators, and schedule the DAG for daily execution.

Full Access

Answer:

Explanation:

Cloud Composer (based on Apache Airflow) is designed for workflow orchestration where complex dependencies exist. To achieve high observability and robustness (specifically the ability to re-run individual steps), the workflow must be decomposed into granular tasks.

Granularity and Re-runs: By defining each step (SFTP fetch, Decrypt, GCS Load, BQ Transform) as a separate task in a DAG, Airflow tracks the state of each individually. If the "BigQuery SQL" step fails, you can "Clear" only that specific task in the UI to re-run it without re-fetching or re-decrypting the data, saving time and costs.

Observability: Each task has its own logs and status in the Airflow UI. Options B and C (Single Operator/Script) are "black boxes"—if they fail, Airflow only knows the entire script failed, making it difficult to pinpoint where and impossible to re-run just the failed portion.

SLA and Notifications: Airflow has built-in sla_miss_callbacks and email_on_failure features. You can set an sla parameter of 7:00 AM to automatically trigger alerts if the process is lagging.

Correcting other options: * A: Splitting the logic between Composer and a separate Scheduler/Dataflow job breaks the end-to-end lineage and makes it harder to manage dependencies and global SLAs.

B & C: As mentioned, using a single operator for multiple logic steps defeats the purpose of an orchestrator's monitoring and retry capabilities.

[Reference: Google Cloud Documentation on Cloud Composer / Airflow:, "An Airflow DAG is a collection of all the tasks you want to run, organized in a way that reflects their relationships and dependencies... Breaking down your workflow into multiple tasks allows for: Individual retries (re-running only failed parts), Parallel execution, and Clearer monitoring in the Airflow web interface." (Source: Key Airflow Concepts), "You can use the SLA (Service Level Agreement) feature in Airflow to track whether a task or DAG takes longer than expected to finish... If a task exceeds its SLA, Airflow can send an email alert or trigger a callback function." (Source: Airflow Documentation - SLAs), , ]

Question # 39

You have data pipelines running on BigQuery, Cloud Dataflow, and Cloud Dataproc. You need to perform health checks and monitor their behavior, and then notify the team managing the pipelines if they fail. You also need to be able to work across multiple projects. Your preference is to use managed products of features of the platform. What should you do?

Export the information to Cloud Stackdriver, and set up an Alerting policy

Run a Virtual Machine in Compute Engine with Airflow, and export the information to Stackdriver

Export the logs to BigQuery, and set up App Engine to read that information and send emails if you find a failure in the logs

Develop an App Engine application to consume logs using GCP API calls, and send emails if you find a failure in the logs

Full Access

Question # 40

Your data science team needs to perform interactive SQL queries on large datasets stored in Apache Parquet format within a Cloud Storage bucket. The team is familiar with Apache Hive and wants to leverage existing HiveQL queries. You need to provide an environment for the team to run their interactive HiveQL queries directly against the data in Cloud Storage. You want to keep operational overhead to a minimum. What should you do?

Load the Parquet data into a BigQuery native table and use the BigQuery Connector for Hive to run the queries.

Install and configure an Apache Hadoop and Hive cluster manually on a group of Compute Engine instances.

Configure BigQuery with an external table definition pointing to the Parquet files.

Deploy a Dataproc cluster with Hive services enabled.

Full Access

Question # 41

Scaling a Cloud Dataproc cluster typically involves ____.

increasing or decreasing the number of worker nodes

increasing or decreasing the number of master nodes

moving memory to run more applications on a single node

deleting applications from unused nodes periodically

Full Access

Question # 42

Google Cloud Bigtable indexes a single value in each row. This value is called the _______.

primary key

unique key

row key

master key

Full Access

Question # 43

Which of these are examples of a value in a sparse vector? (Select 2 answers.)

[0, 5, 0, 0, 0, 0]

[0, 0, 0, 1, 0, 0, 1]

[0, 1]

[1, 0, 0, 0, 0, 0, 0]

Full Access

Question # 44

What are two of the benefits of using denormalized data structures in BigQuery?

Reduces the amount of data processed, reduces the amount of storage required

Increases query speed, makes queries simpler

Reduces the amount of storage required, increases query speed

Reduces the amount of data processed, increases query speed

Full Access

Question # 45

Which Java SDK class can you use to run your Dataflow programs locally?

LocalRunner

DirectPipelineRunner

MachineRunner

LocalPipelineRunner

Full Access

Question # 46

If you're running a performance test that depends upon Cloud Bigtable, all the choices except one below are recommended steps. Which is NOT a recommended step to follow?

Do not use a production instance.

Run your test for at least 10 minutes.

Before you test, run a heavy pre-test for several minutes.

Use at least 300 GB of data.

Full Access

Question # 47

Which of the following job types are supported by Cloud Dataproc (select 3 answers)?

Hive

Pig

YARN

Spark

Full Access

Question # 48

By default, which of the following windowing behavior does Dataflow apply to unbounded data sets?

Windows at every 100 MB of data

Single, Global Window

Windows at every 1 minute

Windows at every 10 minutes

Full Access

Question # 49

You work for a large fast food restaurant chain with over 400,000 employees. You store employee information in Google BigQuery in a Users table consisting of a FirstName field and a LastName field. A member of IT is building an application and asks you to modify the schema and data in BigQuery so the application can query a FullName field consisting of the value of the FirstName field concatenated with a space, followed by the value of the LastName field for each employee. How can you make that data available while minimizing cost?

Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.

Add a new column called FullName to the Users table. Run an UPDATE statement that updates the FullName column for each user with the concatenation of the FirstName and LastName values.

Create a Google Cloud Dataflow job that queries BigQuery for the entire Users table, concatenates the FirstName value and LastName value for each user, and loads the proper values for FirstName, LastName, and FullName into a new table in BigQuery.

Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastName and FullName. Run a BigQuery load job to load the new CSV file into BigQuery.

Full Access

Question # 50

You work for an economic consulting firm that helps companies identify economic trends as they happen. As part of your analysis, you use Google BigQuery to correlate customer data with the average prices of the 100 most common goods sold, including bread, gasoline, milk, and others. The average prices of these goods are updated every 30 minutes. You want to make sure this data stays up to date so you can combine it with other data in BigQuery as cheaply as possible. What should you do?

Load the data every 30 minutes into a new partitioned table in BigQuery.

Store and update the data in a regional Google Cloud Storage bucket and create a federated data source in BigQuery

Store the data in Google Cloud Datastore. Use Google Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Cloud Datastore

Store the data in a file in a regional Google Cloud Storage bucket. Use Cloud Dataflow to query BigQuery and combine the data programmatically with the data stored in Google Cloud Storage.

Full Access

Question # 51

You are designing the database schema for a machine learning-based food ordering service that will predict what users want to eat. Here is some of the information you need to store:

The user profile: What the user likes and doesn’t like to eat

The user account information: Name, address, preferred meal times

The order information: When orders are made, from where, to whom

The database will be used to store all the transactional data of the product. You want to optimize the data schema. Which Google Cloud Platform product should you use?

BigQuery

Cloud SQL

Cloud Bigtable

Cloud Datastore

Full Access

Question # 52

Your company produces 20,000 files every hour. Each data file is formatted as a comma separated values (CSV) file that is less than 4 KB. All files must be ingested on Google Cloud Platform before they can be processed. Your company site has a 200 ms latency to Google Cloud, and your Internet connection bandwidth is limited as 50 Mbps. You currently deploy a secure FTP (SFTP) server on a virtual machine in Google Compute Engine as the data ingestion point. A local SFTP client runs on a dedicated machine to transmit the CSV files as is. The goal is to make reports with data from the previous day available to the executives by 10:00 a.m. each day. This design is barely able to keep up with the current volume, even though the bandwidth utilization is rather low.

You are told that due to seasonality, your company expects the number of files to double for the next three months. Which two actions should you take? (choose two.)

Introduce data compression for each file to increase the rate file of file transfer.

Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.

Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.

Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.

Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer Service to transfer on-premices data to the designated storage bucket.

Full Access

Question # 53

You are choosing a NoSQL database to handle telemetry data submitted from millions of Internet-of-Things (IoT) devices. The volume of data is growing at 100 TB per year, and each data entry has about 100 attributes. The data processing pipeline does not require atomicity, consistency, isolation, and durability (ACID). However, high availability and low latency are required.

You need to analyze the data by querying against individual fields. Which three databases meet your requirements? (Choose three.)

Redis

HBase

MySQL

MongoDB

Cassandra

HDFS with Hive

Full Access

Question # 54

Your company uses a proprietary system to send inventory data every 6 hours to a data ingestion service in the cloud. Transmitted data includes a payload of several fields and the timestamp of the transmission. If there are any concerns about a transmission, the system re-transmits the data. How should you deduplicate the data most efficiency?

Assign global unique identifiers (GUID) to each data entry.

Compute the hash value of each data entry, and compare it with all historical data.

Store each data entry as the primary key in a separate database and apply an index.

Maintain a database table to store the hash value and other metadata for each data entry.

Full Access

Question # 55

Your company is migrating their 30-node Apache Hadoop cluster to the cloud. They want to re-use Hadoop jobs they have already created and minimize the management of the cluster as much as possible. They also want to be able to persist data beyond the life of the cluster. What should you do?

Create a Google Cloud Dataflow job to process the data.

Create a Google Cloud Dataproc cluster that uses persistent disks for HDFS.

Create a Hadoop cluster on Google Compute Engine that uses persistent disks.

Create a Cloud Dataproc cluster that uses the Google Cloud Storage connector.

Create a Hadoop cluster on Google Compute Engine that uses Local SSD disks.

Full Access

Question # 56

Your company is streaming real-time sensor data from their factory floor into Bigtable and they have noticed extremely poor performance. How should the row key be redesigned to improve Bigtable performance on queries that populate real-time dashboards?

Use a row key of the form .

Use a row key of the form #.

Use a row key of the form >##.

Full Access

Question # 57

An external customer provides you with a daily dump of data from their database. The data flows into Google Cloud Storage GCS as comma-separated values (CSV) files. You want to analyze this data in Google BigQuery, but the data could have rows that are formatted incorrectly or corrupted. How should you build this pipeline?

Use federated data sources, and check data in the SQL query.

Enable BigQuery monitoring in Google Stackdriver and create an alert.

Import the data into BigQuery using the gcloud CLI and set max_bad_records to 0.

Run a Google Cloud Dataflow batch pipeline to import the data into BigQuery, and push errors to another dead-letter table for analysis.

Full Access

Question # 58

Your company is using WHILECARD tables to query data across multiple tables with similar names. The SQL statement is currently failing with the following error:

# Syntax error : Expected end of statement but got “-“ at [4:11]

SELECT age

FROM

bigquery-public-data.noaa_gsod.gsod

WHERE

age != 99

AND_TABLE_SUFFIX = ‘1929’

ORDER BY

age DESC

Which table name will make the SQL statement work correctly?

‘bigquery-public-data.noaa_gsod.gsod‘

bigquery-public-data.noaa_gsod.gsod*

‘bigquery-public-data.noaa_gsod.gsod’*

‘bigquery-public-data.noaa_gsod.gsod*`

Full Access

Question # 59

You create an important report for your large team in Google Data Studio 360. The report uses Google BigQuery as its data source. You notice that visualizations are not showing data that is less than 1 hour old. What should you do?

Disable caching by editing the report settings.

Disable caching in BigQuery by editing table details.

Refresh your browser tab showing the visualizations.

Clear your browser history for the past hour then reload the tab showing the virtualizations.

Full Access

Question # 60

You have spent a few days loading data from comma-separated values (CSV) files into the Google BigQuery table CLICK_STREAM. The column DT stores the epoch time of click events. For convenience, you chose a simple schema where every field is treated as the STRING type. Now, you want to compute web session durations of users who visit your site, and you want to change its data type to the TIMESTAMP. You want to minimize the migration effort without making future queries computationally expensive. What should you do?

Delete the table CLICK_STREAM, and then re-create it such that the column DT is of the TIMESTAMP type. Reload the data.

Add a column TS of the TIMESTAMP type to the table CLICK_STREAM, and populate the numeric values from the column TS for each row. Reference the column TS instead of the column DT from now on.

Create a view CLICK_STREAM_V, where strings from the column DT are cast into TIMESTAMP values. Reference the view CLICK_STREAM_V instead of the table CLICK_STREAM from now on.

Add two columns to the table CLICK STREAM: TS of the TIMESTAMP type and IS_NEW of the BOOLEAN type. Reload all data in append mode. For each appended row, set the value of IS_NEW to true. For future queries, reference the column TS instead of the column DT, with the WHERE clause ensuring that the value of IS_NEW must be true.

Construct a query to return every row of the table CLICK_STREAM, while using the built-in function to cast strings from the column DT into TIMESTAMP values. Run the query into a destination table NEW_CLICK_STREAM, in which the column TS is the TIMESTAMP type. Reference the table NEW_CLICK_STREAM instead of the table CLICK_STREAM from now on. In the future, new data is loaded into the table NEW_CLICK_STREAM.

Full Access

Question # 61

MJTelco’s Google Cloud Dataflow pipeline is now ready to start receiving data from the 50,000 installations. You want to allow Cloud Dataflow to scale its compute power up as required. Which Cloud Dataflow pipeline configuration setting should you update?

The zone

The number of workers

The disk size per worker

The maximum number of workers

Full Access

Question # 62

You need to compose visualizations for operations teams with the following requirements:

Which approach meets the requirements?

Load the data into Google Sheets, use formulas to calculate a metric, and use filters/sorting to show only suboptimal links in a table.

Load the data into Google BigQuery tables, write Google Apps Script that queries the data, calculates the metric, and shows only suboptimal rows in a table in Google Sheets.

Load the data into Google Cloud Datastore tables, write a Google App Engine Application that queries all rows, applies a function to derive the metric, and then renders results in a table using the Google charts and visualization API.

Load the data into Google BigQuery tables, write a Google Data Studio 360 report that connects to your data, calculates a metric, and then uses a filter expression to show only suboptimal rows in a table.

Full Access

Question # 63

You need to compose visualization for operations teams with the following requirements:

Telemetry must include data from all 50,000 installations for the most recent 6 weeks (sampling once every minute)

The report must not be more than 3 hours delayed from live data.

The actionable report should only show suboptimal links.

Most suboptimal links should be sorted to the top.

Suboptimal links can be grouped and filtered by regional geography.

User response time to load the report must be <5 seconds.

You create a data source to store the last 6 weeks of data, and create visualizations that allow viewers to see multiple date ranges, distinct geographic regions, and unique installation types. You always show the latest data without any changes to your visualizations. You want to avoid creating and updating new visualizations each month. What should you do?

Look through the current data and compose a series of charts and tables, one for each possiblecombination of criteria.

Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection.

Export the data to a spreadsheet, compose a series of charts and tables, one for each possiblecombination of criteria, and spread them across multiple tabs.

Load the data into relational database tables, write a Google App Engine application that queries all rows, summarizes the data across each criteria, and then renders results using the Google Charts and visualization API.

Full Access

Question # 64

MJTelco is building a custom interface to share data. They have these requirements:

They need to do aggregations over their petabyte-scale datasets.

They need to scan specific time range rows with a very fast response time (milliseconds).

Which combination of Google Cloud Platform products should you recommend?

Cloud Datastore and Cloud Bigtable

Cloud Bigtable and Cloud SQL

BigQuery and Cloud Bigtable

BigQuery and Cloud Storage

Full Access

Question # 65

MJTelco needs you to create a schema in Google Bigtable that will allow for the historical analysis of the last 2 years of records. Each record that comes in is sent every 15 minutes, and contains a unique identifier of the device and a data record. The most common query is for all the data for a given device for a given day. Which schema should you use?

Rowkey: date#device_idColumn data: data_point

Rowkey: dateColumn data: device_id, data_point

Rowkey: device_idColumn data: date, data_point

Rowkey: data_pointColumn data: device_id, date

Rowkey: date#data_pointColumn data: device_id

Full Access

Question # 66

You create a new report for your large team in Google Data Studio 360. The report uses Google BigQuery as its data source. It is company policy to ensure employees can view only the data associated with their region, so you create and populate a table for each region. You need to enforce the regional access policy to the data.

Which two actions should you take? (Choose two.)

Ensure all the tables are included in global dataset.

Ensure each table is included in a dataset for a region.

Adjust the settings for each table to allow a related region-based security group view access.

Adjust the settings for each view to allow a related region-based security group view access.

Adjust the settings for each dataset to allow a related region-based security group view access.

Full Access

Question # 67

Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day’s events. They also want to use streaming ingestion. What should you do?

Create a table called tracking_table and include a DATE column.

Create a partitioned table called tracking_table and include a TIMESTAMP column.

Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.

Create a table called tracking_table with a TIMESTAMP column to represent the day.

Full Access

Summer Sale - Special 70% Discount Offer - Ends in 0d 00h 00m 00s - Coupon code: 70dumps

DumpsTool Header

dumpstool logo

Professional-Data-Engineer Questions and Answers

Answer:

Explanation:

Answer:

Answer:

Answer:

Answer:

Answer:

Explanation:

Answer:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Answer:

Answer:

Explanation:

Answer:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Answer:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Explanation:

Answer:

Answer: