Apache Pig vs. Cloudera Distribution Hadoop (CDH)

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Pig
Score 8.4 out of 10
N/A
Apache Pig is a programming tool for creating MapReduce programs used in Hadoop.N/A
Cloudera Distribution Hadoop (CDH)
Score 4.9 out of 10
N/A
CDH is Cloudera’s 100% open source platform distribution, including Apache Hadoop and built specifically to meet enterprise demands. CDH delivers everything needed for enterprise use right out of the box. By integrating Hadoop with more than a dozen other critical open source projects, Cloudera has created a functionally advanced system that helps you perform end-to-end Big Data workflows.N/A
Pricing
Apache PigCloudera Distribution Hadoop (CDH)
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache PigCloudera Distribution Hadoop (CDH)
Free Trial
NoNo
Free/Freemium Version
YesNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache PigCloudera Distribution Hadoop (CDH)
Considered Both Products
Apache Pig
Chose Apache Pig
Apache Hadoop, Azure Data Lake Storage, Amazon EMR (Elastic MapReduce), Presto (formerly Presto DB), Confluent Platform and Alteryx
Chose Apache Pig
It takes me less time to write a Pig script than get a Spark program running for batch ETL workloads. Compared to Spark, Pig has a steeper learning curve because it employs a proprietary programming language. In one script and one fine, it can handle both Map Reduce and Hadoop. …
Chose Apache Pig
It can accommodate Map Reduce in a single script and a single fine. IT has very much documentation present for easy learning. SQL like queries makes it easy to understand
Chose Apache Pig
Apache Pig might help to start things faster at first and it was one of the best tool years back but it lacks important features that are needed in the data engineering world right now. Pig also has a steeper learning curve since it uses a proprietary language compared to Spark …
Chose Apache Pig
Pig is more focused on scripting in its own PigLatin language rather than integrate into another language like Java/Scala/Python/SQL.
However, for batch ETL workloads, I find that I can write a Pig script quicker than setting up and deploying a Spark program, for example.
Chose Apache Pig
Apache Pig is picked up quickly and can be implemented with very little coding skills. Also the other languages require exact matching of versions during installations which made them somewhat less user-friendly. Also most of the tasks that are done in map reduce can be done …
Chose Apache Pig
I use both Apache Pig and its alternatives like Apache Spark & Apache Hive. Apache Pig was one of the best options in Big Data's initial stages. But now alternatives have taken over the market, rendering Apache Pig behind in the competition. But it is still a better alternative …
Chose Apache Pig
Early on Apache Pig was a great tool for easily writing distributed processing applications without needing to write a complete Java MapReduce job from scratch, but as time as moved on there now better alternatives to get results faster for both ad-hoc analysis and for …
Chose Apache Pig
- Provided better ways for optimized hadoop jobs than Hive but not anymore.
- Spark DSL is much more advanced and compute times are significantly less.
Cloudera Distribution Hadoop (CDH)
Chose Cloudera Distribution Hadoop (CDH)
In terms of functionality there's not much difference, both get the job done. Amazon was more cost-efficient for our team, but this could vary depending on the size of the business. One thing I did notice was that Cloudera seemed to management and spit out our deployments …
Best Alternatives
Apache PigCloudera Distribution Hadoop (CDH)
Small Businesses

No answers on this topic

No answers on this topic

Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 7.1 out of 10
IBM Analytics Engine
IBM Analytics Engine
Score 7.1 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache PigCloudera Distribution Hadoop (CDH)
Likelihood to Recommend
8.2
(0 ratings)
7.0
(0 ratings)
Usability
10.0
(0 ratings)
-
(0 ratings)
Support Rating
6.0
(0 ratings)
-
(0 ratings)
User Testimonials
Apache PigCloudera Distribution Hadoop (CDH)
Likelihood to Recommend
Apache Pig is best suited for ETL-based data processes. It is good in performance in handling and analyzing a large amount of data. it gives faster results than any other similar tool. It is easy to implement and any user with some initial training or some prior SQL knowledge can work on it. Apache Pig is proud to have a large community base globally.
Read full review
Cloudera Distribution Hadoop (CDH) does a lot of things really well - especially on the analytical front. That being said the product is quite expensive. There are seemingly numerous applications that do the same thing on the functional level that are much more cost effecient for enterprise teams. If I were recommending this to a colleague I would let them know the product will absolutely be able to get the job done for their use case, but there are more efficient options
Read full review
Pros
  • Iterative Development - you can write aliases/variables, which are not immediately executed and these are stored in a DAG, which is only evaluated upon dumping or storing another alias.
  • Fast execution - Works with MapReduce, Tez, or Spark execution frameworks to provide fast run times at large scales.
  • Local and remote interoperability - Scripts that depend on testing a small dataset locally before moving to the full thing can simply be done with "pig -x local."
Read full review
  • Solid and robust set of integrations
  • Easy to use and easy to deploy across the enterprise
  • Reliability - never lost any info
  • Simple and clean interface
Read full review
Cons
  • May not fit every need and a SQL-like abstraction may be more effective for some tasks (look at Spark-SQL, Hive, or even an actual DBMS)
  • All Pig jobs are written in a Domain Specific Language so not a lot of transferable knowledge
  • Writing your own User Defined Functions (UDFS) is a nice feature but can be painful to implement in practice
Read full review
  • The price is quite high competitively speaking
  • Hard to learn more robust functions and custom options without experience
Read full review
Usability
It is quick, fast and easy to implement Apache Pig which makes is quite popular to be used.
Read full review
No answers on this topic
Support Rating
The documentation is adequate. I'm not sure how large of an external community there is for support.
Read full review
No answers on this topic
Alternatives Considered
It takes me less time to write a Pig script than get a Spark program running for batch ETL workloads. Compared to Spark, Pig has a steeper learning curve because it employs a proprietary programming language. In one script and one fine, it can handle both Map Reduce and Hadoop. It has a large amount of documentation available to make learning more convenient.
Read full review
In terms of functionality there's not much difference, both get the job done. Amazon was more cost-efficient for our team, but this could vary depending on the size of the business. One thing I did notice was that Cloudera seemed to management and spit out our deployments faster than AWS.
Read full review
Return on Investment
  • Return on Investments are significant considering what it can do with traditional analysis techniques. But, other alternatives like Apache Spark, Hive being more efficient, it is hard to stick to Apache Pig.
  • It can handle large datasets pretty easily compared to SQL. But, again, alternatives are more efficient.
  • While working on unstructured, decentralized dataset, Pig is highly beneficial, as it is not a complete deviation from SQL, but it does not take you in complexity MapReduce as well.
Read full review
  • Saves time by automating typically manual processes (data management, lifecyle AI etc)
  • Quick deployments and analytics allow for faster time-to-value
Read full review
ScreenShots