AWS Glue is a managed extract, transform, and load (ETL) service designed to make it easy for customers to prepare and load data for analytics. With it, users can create and run an ETL job in the AWS Management Console. Users point AWS Glue to data stored on AWS, and AWS Glue discovers data and stores the associated metadata (e.g. table definition and schema) in the AWS Glue Data Catalog. Once cataloged, data is immediately searchable, queryable, and available for ETL.
$0.44
billed per second, 1 minute minimum
RapidMiner
Score 8.9 out of 10
N/A
RapidMiner is a data science and data mining platform, from Altair since the late 2022 acquisition. RapidMiner offers full automation for non-coding domain experts, an integrated JupyterLab environment for seasoned data scientists, and a visual drag-and-drop designer. RapidMiner’s project-based framework helps to ensure that others can build off their work using visual workflows or automated data science.
$7,500
Per User Per Month
Pricing
AWS Glue
RapidMiner
Editions & Modules
per DPU-Hour
$0.44
billed per second, 1 minute minimum
Professional
$7,500.00
Per User Per Month
Enterprise
$15,000.00
Per User Per Month
AI Hub
$54,000.00
Per User Per Month
Offerings
Pricing Offerings
AWS Glue
RapidMiner
Free Trial
No
No
Free/Freemium Version
No
No
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
No setup fee
Additional Details
—
—
More Pricing Information
Community Pulse
AWS Glue
RapidMiner
Features
AWS Glue
RapidMiner
Platform Connectivity
Comparison of Platform Connectivity features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.5
Ratings
13% above category average
Connect to Multiple Data Sources
00 Ratings
10.00 Ratings
Extend Existing Data Sources
00 Ratings
10.00 Ratings
Automatic Data Format Detection
00 Ratings
9.00 Ratings
MDM Integration
00 Ratings
9.00 Ratings
Data Exploration
Comparison of Data Exploration features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.0
Ratings
7% above category average
Visualization
00 Ratings
9.00 Ratings
Interactive Data Analysis
00 Ratings
9.00 Ratings
Data Preparation
Comparison of Data Preparation features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
8.8
Ratings
8% above category average
Interactive Data Cleaning and Enrichment
00 Ratings
9.00 Ratings
Data Transformations
00 Ratings
7.00 Ratings
Data Encryption
00 Ratings
9.00 Ratings
Built-in Processors
00 Ratings
10.00 Ratings
Platform Data Modeling
Comparison of Platform Data Modeling features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.0
Ratings
7% above category average
Multiple Model Development Languages and Tools
00 Ratings
9.00 Ratings
Automated Machine Learning
00 Ratings
9.00 Ratings
Single platform for multiple model development
00 Ratings
9.00 Ratings
Self-Service Model Delivery
00 Ratings
9.00 Ratings
Model Deployment
Comparison of Model Deployment features of Product A and Product B
When the data which requires ETL has different formats, schema, and volume, this service suits them best. So, when the volume is not consistent (typical use-case of healthcare and online shopping), AWS Glue can be the prime choice. When the data is available in both batch and streaming mode, the developer needs to generate a separate codebase. This increases the source code management efforts. So, prefer to go with Glue when the nature of the data is the same (either batched or streamed).
RapidMiner is the best tool to build models on textual data. It is rich in ML algorithms and reduces the need to manually tune the parameters. It automatically optimizes them, thus providing a better solution. RapidMiner again extends great capability for data preparation, its insane connections to almost every data source pulls in the data easily into one environment. And it can comfortably perform data cleaning and process tasks over that. RapidMiner is not so good with image, audio or video data. These data points cannot be used directly in their raw form. They must be transformed into some intermediate form for performing analytics over it. Moreover, there are no connectors to directly pull data from their varied sources. For example, we don't have a connector to read audio data directly from a switch and then convert it to text (although Google speech API is available for audio to text conversion.)
After data cleansing, the team also implemented the best practices for using AWS platform services as a Data Lake, such as job bookmarking for AWS Glue jobs, proper delimiter for the AWS Glue crawlers, partitioning in AWS S3, and transformation to parquet file for compression and faster querying time in Amazon Athena.
Data modernization through combining data from multiple sources into a functioning datasets, rebuilding DW, and resctructuring data sources.
Aims to lessen customer complaints, eliminate manual data extraction requests via SR from different data sources, and Increase accuracy, consistency and speed up reconciliation process.
Wish the tool was more efficient in terms of processing power. The tool takes a lot of CPU processing power, even for a small process on a small data set
Wish there were more options on charts and graphs to visualize the data
Amazon responds in good time once the ticket has been generated but needs to generate tickets frequent because very few sample codes are available, and it's not cover all the scenarios.
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the other advantage.
The other product like RapidMiner Studio that I have used is WEKA. I decided to use RapidMiner because almost all modelling methods and feature selection methods from the Weka machine learning library are available within RapidMiner. Furthermore, RapidMiner Studio is a visual workflow and therefore it is easier to demonstrate and visualise the processes involves in getting the desired results. Visualization of workflow enhances teaching and learning. RapidMiner is rich with algorithms and online learning materials that can assist students in their self-directed learning on data preparation, machine learning, deep learning, text mining, and predictive analytics. Moreover, RapidMiner repository has more than 1500 machine learning algorithms and functions that students can explore for any case study and assignments. The RapidMIner is also an open platform that can seamlessly integrates with other applications programmed with other programming languages like R and Python.