AWS Glue vs. RapidMiner

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
AWS Glue
Score 7.5 out of 10
N/A
AWS Glue is a managed extract, transform, and load (ETL) service designed to make it easy for customers to prepare and load data for analytics. With it, users can create and run an ETL job in the AWS Management Console. Users point AWS Glue to data stored on AWS, and AWS Glue discovers data and stores the associated metadata (e.g. table definition and schema) in the AWS Glue Data Catalog. Once cataloged, data is immediately searchable, queryable, and available for ETL.
$0.44
billed per second, 1 minute minimum
RapidMiner
Score 8.9 out of 10
N/A
RapidMiner is a data science and data mining platform, from Altair since the late 2022 acquisition. RapidMiner offers full automation for non-coding domain experts, an integrated JupyterLab environment for seasoned data scientists, and a visual drag-and-drop designer. RapidMiner’s project-based framework helps to ensure that others can build off their work using visual workflows or automated data science.
$7,500
Per User Per Month
Pricing
AWS GlueRapidMiner
Editions & Modules
per DPU-Hour
$0.44
billed per second, 1 minute minimum
Professional
$7,500.00
Per User Per Month
Enterprise
$15,000.00
Per User Per Month
AI Hub
$54,000.00
Per User Per Month
Offerings
Pricing Offerings
AWS GlueRapidMiner
Free Trial
NoNo
Free/Freemium Version
NoNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
AWS GlueRapidMiner
Considered Both Products
AWS Glue
Chose AWS Glue
Informatica Intelligent Cloud Integration Services and Informatica PowerCenter
Chose AWS Glue
AWS Glue is a fully managed ETL service that automates many ETL tasks, making it easier to set AWS Glue simplifies ETL through a visual interface and automated code generation.
Chose AWS Glue
AWS Glue is easier to use and has more and better features compared to it. And more documentation and tutorials and labs are widely available on the internet about AWS Glue which in turn helps in easier implementation of the spark jobs. Auto scaling is an added advantage. It's …
Chose AWS Glue
The main reason we choose AWS Glue over Talend open studio 1) Does not support Spark 2) Run only on java 3) not really feasible solution for heavy workloads 4) most of the cases need customer support 5) no proper documentation is available
Chose AWS Glue
AWS Glue is a managed service. It was easier for us to integrate it into our stack since we are already an AWS shop. It saved us the headache of managing a 3rd part service.
Chose AWS Glue
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the …
Chose AWS Glue
Glue comes in form of a managed service. However, the AWS data pipeline puts additional responsibility to manage the infrastructure. We were not requiring fine-grained control of the hardware which the AWS data pipeline provides. We also want to park our data on DynamoDB. AWS …
Chose AWS Glue
We are already in AWS services, so AWS glue is the first choice for us. But for the comparison of ETL job making and process time, it's way faster for other services.
Chose AWS Glue
Glue is easier especially if you are already in AWS. It easily integrates to other AWS services. Compliments well with Amazon Athena, S3, and Lake Formation. Compared to Snowflake, it is also much much cheaper and you don't have to build outside AWS. Support is also good if you …
RapidMiner
Chose RapidMiner
We tried different data tools and we figured we give RapidMinder Studio a shot as one of our employees had experience with it, and when compared to some of the other tools that we used it was the best fit among the test group that we used. Overall it was a little more fluid and …
Chose RapidMiner
For me, the best advantage to use RapidMiner is the ease of use to learn and deploy new processes. Yo don't need to code, you learn fast and it's really flexible when it comes to transforming data. Knime is also good, but not so flexible, and visually less attractive. Pentaho …
Chose RapidMiner
I found RapidMiner to be in a class of its own. It's easy to use, yet extremely powerful for full data analysis.
Chose RapidMiner
The other product like RapidMiner Studio that I have used is WEKA. I decided to use RapidMiner because almost all modelling methods and feature selection methods from the Weka machine learning library are available within RapidMiner. Furthermore, RapidMiner Studio is a visual …
Chose RapidMiner
Used R and RapidMiner Studio. The main advantage for RapidMiner Studio is the reduced need to program. It has a much smaller learning curve, and it is easy to start using the tool and analyzing from day one.
Chose RapidMiner
RapidMiner is much easier and faster to use plus it interfaces with databases easily.
Chose RapidMiner
We selected RapidMiner due to ease of use and a comfortable user interface. It stacks up very well against these tools in the predictive analytics space. For basic analytics and data reporting, we chose QlikView and Qlik Sense as a more robust reporting platform.
Chose RapidMiner
SPSS and SAS are too expensive. Their interfaces are excellent, but the price point is quite high making them inappropriate for higher education. KNIME is my second choice tool in this space, but it doesn't have the same long established english-speaking user community as …
Chose RapidMiner
It's a heck of a lot better than Python, i.e., it's much quicker to get results with RapidMiner. And RapidMiner is less error prone that coding.
Chose RapidMiner
The best part about RapidMiner is it mainly focus on machine learning algorithms whereas other tools focus on mainly the extract transform load (ETL) process. It can serve for all the KDD (Knowledge data discovery) process stages e.g. data cleaning, transformation, modeling and …
Chose RapidMiner
RapidMIner Studio is freely available and requires no programming skills. When compared with other free analytics tools, its graphical and analytical capabilities are far superior.
Chose RapidMiner
You simply cannot do everything with RapidMiner, it is just one tool in your arsenal. I like using Python directly much better with tools such as Jupyter Notebook in conjunction with JupyterHub.
Chose RapidMiner
The problem with R was that you had to code everything yourself and it doesn't do that well with large amounts of data. At the same time the advantage it provided was it has a large user base which means that you could get help easily.
Features
AWS GlueRapidMiner
Platform Connectivity
Comparison of Platform Connectivity features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.5
Ratings
13% above category average
Connect to Multiple Data Sources00 Ratings10.00 Ratings
Extend Existing Data Sources00 Ratings10.00 Ratings
Automatic Data Format Detection00 Ratings9.00 Ratings
MDM Integration00 Ratings9.00 Ratings
Data Exploration
Comparison of Data Exploration features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.0
Ratings
7% above category average
Visualization00 Ratings9.00 Ratings
Interactive Data Analysis00 Ratings9.00 Ratings
Data Preparation
Comparison of Data Preparation features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
8.8
Ratings
8% above category average
Interactive Data Cleaning and Enrichment00 Ratings9.00 Ratings
Data Transformations00 Ratings7.00 Ratings
Data Encryption00 Ratings9.00 Ratings
Built-in Processors00 Ratings10.00 Ratings
Platform Data Modeling
Comparison of Platform Data Modeling features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.0
Ratings
7% above category average
Multiple Model Development Languages and Tools00 Ratings9.00 Ratings
Automated Machine Learning00 Ratings9.00 Ratings
Single platform for multiple model development00 Ratings9.00 Ratings
Self-Service Model Delivery00 Ratings9.00 Ratings
Model Deployment
Comparison of Model Deployment features of Product A and Product B
AWS Glue
-
Ratings
RapidMiner
9.0
Ratings
5% above category average
Flexible Model Publishing Options00 Ratings9.00 Ratings
Security, Governance, and Cost Controls00 Ratings9.00 Ratings
Best Alternatives
AWS GlueRapidMiner
Small Businesses
IBM SPSS Modeler
IBM SPSS Modeler
Score 7.1 out of 10
Jupyter Notebook
Jupyter Notebook
Score 9.4 out of 10
Medium-sized Companies
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
Posit
Posit
Score 10.0 out of 10
Enterprises
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
Posit
Posit
Score 10.0 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
AWS GlueRapidMiner
Likelihood to Recommend
7.0
(0 ratings)
10.0
(0 ratings)
Likelihood to Renew
-
(0 ratings)
9.0
(0 ratings)
Usability
7.0
(0 ratings)
9.0
(0 ratings)
Support Rating
7.0
(0 ratings)
-
(0 ratings)
User Testimonials
AWS GlueRapidMiner
Likelihood to Recommend
When the data which requires ETL has different formats, schema, and volume, this service suits them best. So, when the volume is not consistent (typical use-case of healthcare and online shopping), AWS Glue can be the prime choice. When the data is available in both batch and streaming mode, the developer needs to generate a separate codebase. This increases the source code management efforts. So, prefer to go with Glue when the nature of the data is the same (either batched or streamed).
Read full review
RapidMiner is the best tool to build models on textual data. It is rich in ML algorithms and reduces the need to manually tune the parameters. It automatically optimizes them, thus providing a better solution. RapidMiner again extends great capability for data preparation, its insane connections to almost every data source pulls in the data easily into one environment. And it can comfortably perform data cleaning and process tasks over that. RapidMiner is not so good with image, audio or video data. These data points cannot be used directly in their raw form. They must be transformed into some intermediate form for performing analytics over it. Moreover, there are no connectors to directly pull data from their varied sources. For example, we don't have a connector to read audio data directly from a switch and then convert it to text (although Google speech API is available for audio to text conversion.)
Read full review
Pros
  • After data cleansing, the team also implemented the best practices for using AWS platform services as a Data Lake, such as job bookmarking for AWS Glue jobs, proper delimiter for the AWS Glue crawlers, partitioning in AWS S3, and transformation to parquet file for compression and faster querying time in Amazon Athena.
  • Data modernization through combining data from multiple sources into a functioning datasets, rebuilding DW, and resctructuring data sources.
  • Aims to lessen customer complaints, eliminate manual data extraction requests via SR from different data sources, and Increase accuracy, consistency and speed up reconciliation process.
Read full review
  • RapidMiner Studio offers a superb user interface with an intuitive workflow paradigm that is very easy to learn.
  • RapidMiner Studio’s operators make it a complete and powerful tool for data preprocessing, data visualization, and data mining/analytics.
  • RapidMiner Studio provides excellent documentation, countless worked examples, training and support via a large user community.
  • Every problem is solved using a sequence of operators.
  • Statistical analysis capabilities offered with the T-Test, ANOVA, Grouped ANOVA, and ANOVA Matrix operators.
  • Textual data mining operators.
  • Web-based and cloud computing capabilities.
  • Visualization capabilities.
  • Marketplace Extensions – especially Finance And Economics.
  • Process portability.
Read full review
Cons
  • It’s integration with other cloud vendors is bit difficult
  • If it can support non SQL based databases as well, it would be powerful.
  • Real time data synchronisation in data source is missing
Read full review
  • Wish the tool was more efficient in terms of processing power. The tool takes a lot of CPU processing power, even for a small process on a small data set
  • Wish there were more options on charts and graphs to visualize the data
Read full review
Likelihood to Renew
No answers on this topic
Very fast and user-friendly tool
Read full review
Usability
I personally found it very usable for a data engineer's day job, particularly for performing ETL and managing the data pipelines.
Read full review
Very use to use and learn
Read full review
Support Rating
Amazon responds in good time once the ticket has been generated but needs to generate tickets frequent because very few sample codes are available, and it's not cover all the scenarios.
Read full review
No answers on this topic
Alternatives Considered
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the other advantage.
Read full review
The other product like RapidMiner Studio that I have used is WEKA. I decided to use RapidMiner because almost all modelling methods and feature selection methods from the Weka machine learning library are available within RapidMiner. Furthermore, RapidMiner Studio is a visual workflow and therefore it is easier to demonstrate and visualise the processes involves in getting the desired results. Visualization of workflow enhances teaching and learning. RapidMiner is rich with algorithms and online learning materials that can assist students in their self-directed learning on data preparation, machine learning, deep learning, text mining, and predictive analytics. Moreover, RapidMiner repository has more than 1500 machine learning algorithms and functions that students can explore for any case study and assignments. The RapidMIner is also an open platform that can seamlessly integrates with other applications programmed with other programming languages like R and Python.
Read full review
Return on Investment
  • Positive Impact :- after ETL we can able to do some kind of automation
  • Negative :- At some point of time it can hamper the cost but not really
Read full review
  • We saved over $100k on our direct mail program by not mailing to those unlikely to respond to our mailings based on our predictive analysis.
  • Our CX team has saved countless hours by automating call scripts to isolate key phrases and code each call.
Read full review
ScreenShots