Apache Iceberg vs. AWS Glue

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Iceberg
Score 7.9 out of 10
N/A
Apache Iceberg is an open table format for huge analytic datasets. Iceberg adds tables to compute engines including Spark, Trino, PrestoDB, Flink, Hive and Impala using a high-performance table format that works just like a SQL table. Iceberg is used in production where a single table can contain tens of petabytes of data and even these huge tables can be read without a distributed SQL engine. It was designed to solve correctness problems in eventually-consistent cloud object stores.N/A
AWS Glue
Score 7.5 out of 10
N/A
AWS Glue is a managed extract, transform, and load (ETL) service designed to make it easy for customers to prepare and load data for analytics. With it, users can create and run an ETL job in the AWS Management Console. Users point AWS Glue to data stored on AWS, and AWS Glue discovers data and stores the associated metadata (e.g. table definition and schema) in the AWS Glue Data Catalog. Once cataloged, data is immediately searchable, queryable, and available for ETL.
$0.44
billed per second, 1 minute minimum
Pricing
Apache IcebergAWS Glue
Editions & Modules
No answers on this topic
per DPU-Hour
$0.44
billed per second, 1 minute minimum
Offerings
Pricing Offerings
Apache IcebergAWS Glue
Free Trial
NoNo
Free/Freemium Version
YesNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache IcebergAWS Glue
Considered Both Products
Apache Iceberg

No answer on this topic

AWS Glue
Chose AWS Glue
Informatica Intelligent Cloud Integration Services and Informatica PowerCenter
Chose AWS Glue
AWS Glue is a fully managed ETL service that automates many ETL tasks, making it easier to set AWS Glue simplifies ETL through a visual interface and automated code generation.
Chose AWS Glue
AWS Glue is easier to use and has more and better features compared to it. And more documentation and tutorials and labs are widely available on the internet about AWS Glue which in turn helps in easier implementation of the spark jobs. Auto scaling is an added advantage. It's …
Chose AWS Glue
The main reason we choose AWS Glue over Talend open studio 1) Does not support Spark 2) Run only on java 3) not really feasible solution for heavy workloads 4) most of the cases need customer support 5) no proper documentation is available
Chose AWS Glue
AWS Glue is a managed service. It was easier for us to integrate it into our stack since we are already an AWS shop. It saved us the headache of managing a 3rd part service.
Chose AWS Glue
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the …
Chose AWS Glue
Glue comes in form of a managed service. However, the AWS data pipeline puts additional responsibility to manage the infrastructure. We were not requiring fine-grained control of the hardware which the AWS data pipeline provides. We also want to park our data on DynamoDB. AWS …
Chose AWS Glue
We are already in AWS services, so AWS glue is the first choice for us. But for the comparison of ETL job making and process time, it's way faster for other services.
Chose AWS Glue
Glue is easier especially if you are already in AWS. It easily integrates to other AWS services. Compliments well with Amazon Athena, S3, and Lake Formation. Compared to Snowflake, it is also much much cheaper and you don't have to build outside AWS. Support is also good if you …
Best Alternatives
Apache IcebergAWS Glue
Small Businesses
DBeaver
DBeaver
Score 9.2 out of 10
IBM SPSS Modeler
IBM SPSS Modeler
Score 7.1 out of 10
Medium-sized Companies
DBeaver
DBeaver
Score 9.2 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
Enterprises
DBeaver
DBeaver
Score 9.2 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache IcebergAWS Glue
Likelihood to Recommend
-
(0 ratings)
7.0
(0 ratings)
Usability
-
(0 ratings)
7.0
(0 ratings)
Support Rating
-
(0 ratings)
7.0
(0 ratings)
User Testimonials
Apache IcebergAWS Glue
Likelihood to Recommend
No answers on this topic
When the data which requires ETL has different formats, schema, and volume, this service suits them best. So, when the volume is not consistent (typical use-case of healthcare and online shopping), AWS Glue can be the prime choice. When the data is available in both batch and streaming mode, the developer needs to generate a separate codebase. This increases the source code management efforts. So, prefer to go with Glue when the nature of the data is the same (either batched or streamed).
Read full review
Pros
No answers on this topic
  • After data cleansing, the team also implemented the best practices for using AWS platform services as a Data Lake, such as job bookmarking for AWS Glue jobs, proper delimiter for the AWS Glue crawlers, partitioning in AWS S3, and transformation to parquet file for compression and faster querying time in Amazon Athena.
  • Data modernization through combining data from multiple sources into a functioning datasets, rebuilding DW, and resctructuring data sources.
  • Aims to lessen customer complaints, eliminate manual data extraction requests via SR from different data sources, and Increase accuracy, consistency and speed up reconciliation process.
Read full review
Cons
No answers on this topic
  • It’s integration with other cloud vendors is bit difficult
  • If it can support non SQL based databases as well, it would be powerful.
  • Real time data synchronisation in data source is missing
Read full review
Usability
No answers on this topic
I personally found it very usable for a data engineer's day job, particularly for performing ETL and managing the data pipelines.
Read full review
Support Rating
No answers on this topic
Amazon responds in good time once the ticket has been generated but needs to generate tickets frequent because very few sample codes are available, and it's not cover all the scenarios.
Read full review
Alternatives Considered
No answers on this topic
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the other advantage.
Read full review
Return on Investment
No answers on this topic
  • Positive Impact :- after ETL we can able to do some kind of automation
  • Negative :- At some point of time it can hamper the cost but not really
Read full review
ScreenShots