Apache Pig vs. Azure Data Lake Storage

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Pig
Score 8.4 out of 10
N/A
Apache Pig is a programming tool for creating MapReduce programs used in Hadoop.N/A
Azure Data Lake Storage
Score 9.6 out of 10
N/A
Azure Data Lake Storage Gen2 is a highly scalable and cost-effective data lake solution for big data analytics. It combines the power of a high-performance file system with massive scale and economy to help you speed your time to insight. Data Lake Storage Gen2 extends Azure Blob Storage capabilities and is optimized for analytics workloads.N/A
Pricing
Apache PigAzure Data Lake Storage
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache PigAzure Data Lake Storage
Free Trial
NoNo
Free/Freemium Version
YesNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache PigAzure Data Lake Storage
Considered Both Products
Apache Pig
Chose Apache Pig
Apache Hadoop, Azure Data Lake Storage, Amazon EMR (Elastic MapReduce), Presto (formerly Presto DB), Confluent Platform and Alteryx
Chose Apache Pig
It takes me less time to write a Pig script than get a Spark program running for batch ETL workloads. Compared to Spark, Pig has a steeper learning curve because it employs a proprietary programming language. In one script and one fine, it can handle both Map Reduce and Hadoop. …
Chose Apache Pig
It can accommodate Map Reduce in a single script and a single fine. IT has very much documentation present for easy learning. SQL like queries makes it easy to understand
Chose Apache Pig
Apache Pig might help to start things faster at first and it was one of the best tool years back but it lacks important features that are needed in the data engineering world right now. Pig also has a steeper learning curve since it uses a proprietary language compared to Spark …
Chose Apache Pig
Pig is more focused on scripting in its own PigLatin language rather than integrate into another language like Java/Scala/Python/SQL.
However, for batch ETL workloads, I find that I can write a Pig script quicker than setting up and deploying a Spark program, for example.
Chose Apache Pig
Apache Pig is picked up quickly and can be implemented with very little coding skills. Also the other languages require exact matching of versions during installations which made them somewhat less user-friendly. Also most of the tasks that are done in map reduce can be done …
Chose Apache Pig
I use both Apache Pig and its alternatives like Apache Spark & Apache Hive. Apache Pig was one of the best options in Big Data's initial stages. But now alternatives have taken over the market, rendering Apache Pig behind in the competition. But it is still a better alternative …
Chose Apache Pig
Early on Apache Pig was a great tool for easily writing distributed processing applications without needing to write a complete Java MapReduce job from scratch, but as time as moved on there now better alternatives to get results faster for both ad-hoc analysis and for …
Chose Apache Pig
- Provided better ways for optimized hadoop jobs than Hive but not anymore.
- Spark DSL is much more advanced and compute times are significantly less.
Azure Data Lake Storage
Chose Azure Data Lake Storage
We have used both Hadoop and GCS buckets for our storage needs of very large healthcare data. In terms of comparison with the Hadoop distributed Files system, Azure Data Lake Storage always stands in a far better position due to easy integration with various latest and widely …
Chose Azure Data Lake Storage
Azure Data Lake Storage from a functionality perspective is a much easier solution to work with. It's implementation from Amazon EMR went smooth, and continued usage is definitely better. However, Amazon EMR was significantly cheaper overall between the high transaction fees …
Chose Azure Data Lake Storage
We chose Azure Data Lake due to the fact that it was already a product under the Azure application suite. We didn't have to focus on integrating another 3rd party application within our environment. Also due to the fact Azure Data Lake scales its storage pools very efficiently, …
Chose Azure Data Lake Storage
We decided long ago to develop for the Azure platform, so we only evaluate products from within Azure. And Azure Data Lake Storage is really the dominant offering within its space. But to give you a comparison, previously we used to use Azure SQL Database for our analytical …
Chose Azure Data Lake Storage
Microsoft solutions provide great harmony in end-to-end data value creation, and Azure Data Lake Storage is highly compatible with other analytical solutions, e.g., Azure Data Factory and Databricks. So I would say that it is at the heart of the analytical solution in the …
Chose Azure Data Lake Storage
Snowflake, Apache Hive, Google BigQuery, Alteryx and Databricks Lakehouse Platform (Unified Analytics Platform)
Chose Azure Data Lake Storage
AWS charges you on an hourly basis but Azure has a pricing model of per minute charge. In terms of short term subscriptions, Azure has more flexibility but it is more expensive. Azure has a much better hybrid cloud support in comparison with AWS. AWS provides direct connections …
Chose Azure Data Lake Storage
I am much more familiar with Snowflake. I thought it was fairly straightforward to use and did not have to learn much syntax. With Azure Data Lake Storage, I have had to learn some new syntax and thought there was a steeper learning curve. We selected it because of cost savings.
Chose Azure Data Lake Storage
Simpler to use, in my opinion. It is also slightly cheaper.
Chose Azure Data Lake Storage
We looked at the Amazon solution and it did not play as well with our existing tools and added a layer of maintenance that we were not willing to take on at the time. We thought that our Microsoft contract and support were good and that our internal team had the knowledge to …
Chose Azure Data Lake Storage
The Azure Data Lake solution is designed for organizations that want to take advantage of big data. It provides a data platform that can help developers, data scientists, and analysts store data of any size and format and perform all types of processing and analytics across …
Best Alternatives
Apache PigAzure Data Lake Storage
Small Businesses

No answers on this topic

Amazon S3 Glacier
Amazon S3 Glacier
Score 9.1 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Azure Blob Storage
Azure Blob Storage
Score 9.7 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 7.1 out of 10
Azure Blob Storage
Azure Blob Storage
Score 9.7 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache PigAzure Data Lake Storage
Likelihood to Recommend
8.2
(0 ratings)
8.2
(0 ratings)
Usability
10.0
(0 ratings)
-
(0 ratings)
Support Rating
6.0
(0 ratings)
-
(0 ratings)
User Testimonials
Apache PigAzure Data Lake Storage
Likelihood to Recommend
Apache Pig is best suited for ETL-based data processes. It is good in performance in handling and analyzing a large amount of data. it gives faster results than any other similar tool. It is easy to implement and any user with some initial training or some prior SQL knowledge can work on it. Apache Pig is proud to have a large community base globally.
Read full review
Azure Data Lake storage is well suited for applications/use cases within organizations where capturing and storing large amounts of data in any format is required, primarily for storing and processing purposes. It's an easy and cost-effective cloud solution for your application data. The ability to integrate with other Azure Services like Azure Databricks and Azure Data Factory is superb.
Read full review
Pros
  • Iterative Development - you can write aliases/variables, which are not immediately executed and these are stored in a DAG, which is only evaluated upon dumping or storing another alias.
  • Fast execution - Works with MapReduce, Tez, or Spark execution frameworks to provide fast run times at large scales.
  • Local and remote interoperability - Scripts that depend on testing a small dataset locally before moving to the full thing can simply be done with "pig -x local."
Read full review
  • Azure Data Lake Storage is extremely scalable. It allows us to scale up or down endlessly based on what we need including replication.
  • In terms of security, Azure Data Lake Storage fits our requirements really well as we can monitor and encrypt seamlessly. We can also assign permissions through roles and grant network-level access.
  • Due to the fact that it can scale, we are able to monitor the cost of storage and any given time and make financial decisions about our infrastructure based on how small or big we want to scale.
Read full review
Cons
  • May not fit every need and a SQL-like abstraction may be more effective for some tasks (look at Spark-SQL, Hive, or even an actual DBMS)
  • All Pig jobs are written in a Domain Specific Language so not a lot of transferable knowledge
  • Writing your own User Defined Functions (UDFS) is a nice feature but can be painful to implement in practice
Read full review
  • I'd like to see a better cross-platform native client. Azure Data Explorer is fine, but it's far from the "SSMS" kind of experience SQL Server users are used to.
  • Listing a large number of file is somewhat problematic and slow. Using the native C# library, running directly on an Azure VM, it can take several hours to list just a couple million files.
  • Switching from V1 to V2 requires the creation of a new Storage Account and that's pretty inconvenient.
Read full review
Usability
It is quick, fast and easy to implement Apache Pig which makes is quite popular to be used.
Read full review
No answers on this topic
Support Rating
The documentation is adequate. I'm not sure how large of an external community there is for support.
Read full review
No answers on this topic
Alternatives Considered
It takes me less time to write a Pig script than get a Spark program running for batch ETL workloads. Compared to Spark, Pig has a steeper learning curve because it employs a proprietary programming language. In one script and one fine, it can handle both Map Reduce and Hadoop. It has a large amount of documentation available to make learning more convenient.
Read full review
The Azure Data Lake solution is designed for organizations that want to take advantage of big data. It provides a data platform that can help developers, data scientists, and analysts store data of any size and format and perform all types of processing and analytics across multiple platforms and programming languages. It can work with your existing solutions, such as identity management and security solutions. It also integrates with other data warehouses and cloud environments. It can be useful for organizations that need the above softwares.
Read full review
Return on Investment
  • Return on Investments are significant considering what it can do with traditional analysis techniques. But, other alternatives like Apache Spark, Hive being more efficient, it is hard to stick to Apache Pig.
  • It can handle large datasets pretty easily compared to SQL. But, again, alternatives are more efficient.
  • While working on unstructured, decentralized dataset, Pig is highly beneficial, as it is not a complete deviation from SQL, but it does not take you in complexity MapReduce as well.
Read full review
  • The cost can be high for more advanced work. In some cases, for instance, time limits and lab runtimes may be too short if you are too slow to learn what is explained as you go along.
  • promote flexible team communication. You can create different spaces for different teams, and share files and tasks.
Read full review
ScreenShots