Hadoop is an open source software from Apache, supporting distributed processing and data storage. Hadoop is popular for its scalability, reliability, and functionality available across commoditized hardware.
N/A
Azure Data Lake Storage
Score 9.6 out of 10
N/A
Azure Data Lake Storage Gen2 is a highly scalable and cost-effective data lake solution for big data analytics. It combines the power of a high-performance file system with massive scale and economy to help you speed your time to insight. Data Lake Storage Gen2 extends Azure Blob Storage capabilities and is optimized for analytics workloads.
N/A
Pricing
Apache Hadoop
Azure Data Lake Storage
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Hadoop
Azure Data Lake Storage
Free Trial
No
No
Free/Freemium Version
Yes
No
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
No setup fee
Additional Details
—
—
More Pricing Information
Community Pulse
Apache Hadoop
Azure Data Lake Storage
Considered Both Products
Hadoop
Verified User
Anonymous
Chose Hadoop
It’s open source nature it’s community support its being configurable
Different departments of my organization have been getting the benefit from Apache Hadoop as it serves the purpose of saving lives when large amounts of data is unable to be converted and processed in a timely manner from a node or a simple computer. Hadoop also has an easier …
I feel that this is a highly reliable and scalable solution computing technology that is highly capable of processing large data sets across multiple servers and thousands of machines in a well-defined and distributed manner. Apache Hadoop can automatically scale up the number …
Spark is a good alternative to Hadoop that can have faster querying and processing performance and can offer more flexibility in terms of applications that it can support.
Google Bigquery has also been a great alternative and is especially great in terms of ease of use. The …
MariaDB - Better to be already in the cloud you will use it for. Issues have improved as it has matured over the year.s CockroachDB - Not nearly as performant (even out of the box) as Apache Hadoop. More configurations required just to make it work. In memory cacheing is an issue.
Hadoop utilizes a SQL structure, which is great. You pay less for the services, but it's definitely less of an enterprise-level option and more just a good place to store your seldom-used data. Teradata and AWS are a lot faster in returning queries than Hadoop, but you pay …
Vice President, Chief Architect, Development Manager and Software Engineer
Chose Hadoop
Hands down, Hadoop is less expensive than the other platforms we considered. Cloudera was easier to set up but the expense ruled it out. MS-SQL didn't have the performance we saw with the Hadoop clusters and was more expensive. We considered MS-SQL mainly for its ability …
When comparing to the sophistication of IBM GPFS (Spectrum Scale) to Hadoop, it is clear that Spectrum Scale is a much better choice. That is maybe something you don't want to hear, but in all of our research, this has been the final decision of the client.
Apache Spark can be considered as an alternative because of its similar capabilities around processing and storing big data. The reason we went with Hadoop was the literature available online and integration capability with platforms like R Studio. The popularity of Hadoop has …
Hadoop offers a scalable, cost-effective and highly available solution for big data storage and processing. The use of a non-proprietary physical layer greatly reduces dependency on technology. It also offers elastic dimensioning capability when deployed on virtual machines or …
I haven't worked with other Big Data aggregation services like Hadoop. As far as I know, Hadoop is the leading choice in this field with good cause. There is a lot of community support, custom modules, paid consultants, free and paid training. All this makes it an ideal choice …
No SQL database were evaluated along with MPP platform. Hadoop performs very well compared to the other platforms. Also since lot of investment goes into Hadoop there is a good chance of getting what one needs from the developer community.
As I am new to the hadoop ecosystem I have not used or evaluated any other similar products at this time. This was handed to me from a previous much older installation that was very under utilized. Our new platform will be working the new cluster much harder with jobs that run …
Hadoop was a cheaper alternative to Amazon. Since I had to pay for every minute I use with Amazon, I had to make sure multiple times that the code was good enough before I purchased with Amazon. But since Hadoop was available on the cluster, I had the opportunity to code on the …
Hadoop being open source, is cheaper to use and do POCs for clients. Cloudera, Hortonworks and MapR also compete to contribute to open source Hadoop and keep their product conceptually similar to Hadoop.
Apache Spark has an in memory processing model, making it powerful for lightning fast data processing. Apache Spark also exposes Scala and Python in APIs which is one of the most commonly used programming languages in data analytic and data processing domains.
Not used any other product than Hadoop and I don't think our company will switch to any other product, as Hadoop is providing excellent results. Our company is growing rapidly, Hadoop helps to keep up our performance and meet customer expectations. We also use HDFS which …
Hadoop provides storage for large data sets and a powerful processing model to crunch and transform huge amounts of data. It does not assume the underlying hardware or infrastructure and enables the users to build data processing infrastructure from commodity hardware. All the …
Processing of big data has been the ultimate need for the me choosing Hadoop. Big data is massive and messy, and it’s coming at you uncontrolled. Data are gathered to be analyzed to discover patterns and correlations that could not be initially apparent, but might be useful in …
Hadoop solves lot of problems (involving unstructured data and huge volumes of data ) better than traditional database systems . And it is completely free and open source ( so lots of cost savings ). Data analysis is very fast when compared to old systems, resulting in more …
We have used both Hadoop and GCS buckets for our storage needs of very large healthcare data. In terms of comparison with the Hadoop distributed Files system, Azure Data Lake Storage always stands in a far better position due to easy integration with various latest and widely …
Azure Data Lake Storage from a functionality perspective is a much easier solution to work with. It's implementation from Amazon EMR went smooth, and continued usage is definitely better. However, Amazon EMR was significantly cheaper overall between the high transaction fees …
We chose Azure Data Lake due to the fact that it was already a product under the Azure application suite. We didn't have to focus on integrating another 3rd party application within our environment. Also due to the fact Azure Data Lake scales its storage pools very efficiently, …
We decided long ago to develop for the Azure platform, so we only evaluate products from within Azure. And Azure Data Lake Storage is really the dominant offering within its space. But to give you a comparison, previously we used to use Azure SQL Database for our analytical …
Microsoft solutions provide great harmony in end-to-end data value creation, and Azure Data Lake Storage is highly compatible with other analytical solutions, e.g., Azure Data Factory and Databricks. So I would say that it is at the heart of the analytical solution in the …
AWS charges you on an hourly basis but Azure has a pricing model of per minute charge. In terms of short term subscriptions, Azure has more flexibility but it is more expensive. Azure has a much better hybrid cloud support in comparison with AWS. AWS provides direct connections …
I am much more familiar with Snowflake. I thought it was fairly straightforward to use and did not have to learn much syntax. With Azure Data Lake Storage, I have had to learn some new syntax and thought there was a steeper learning curve. We selected it because of cost savings.
We looked at the Amazon solution and it did not play as well with our existing tools and added a layer of maintenance that we were not willing to take on at the time. We thought that our Microsoft contract and support were good and that our internal team had the knowledge to …
The Azure Data Lake solution is designed for organizations that want to take advantage of big data. It provides a data platform that can help developers, data scientists, and analysts store data of any size and format and perform all types of processing and analytics across …
Apache Hadoop (and its subsequent add-ons) are well-suited to larger, unstructured data flows, such as aggregation of web traffic or advertising. Geospatial algorithms and their outputs are well-suited for this kind of aggregation as structuring that data is challenging, but leaving it unstructured and performing queries as-needed is a better fit for most business models. With the advent of data science, I would expect Hadoop fits a LOT of their initial outputs quite well.
Azure Data Lake storage is well suited for applications/use cases within organizations where capturing and storing large amounts of data in any format is required, primarily for storing and processing purposes. It's an easy and cost-effective cloud solution for your application data. The ability to integrate with other Azure Services like Azure Databricks and Azure Data Factory is superb.
Azure Data Lake Storage is extremely scalable. It allows us to scale up or down endlessly based on what we need including replication.
In terms of security, Azure Data Lake Storage fits our requirements really well as we can monitor and encrypt seamlessly. We can also assign permissions through roles and grant network-level access.
Due to the fact that it can scale, we are able to monitor the cost of storage and any given time and make financial decisions about our infrastructure based on how small or big we want to scale.
Hadoop is a batch oriented processing framework, it lacks real time or stream processing.
Hadoop's HDFS file system is not a POSIX compliant file system and does not work well with small files, especially smaller than the default block size.
Hadoop cannot be used for running interactive jobs or analytics.
I'd like to see a better cross-platform native client. Azure Data Explorer is fine, but it's far from the "SSMS" kind of experience SQL Server users are used to.
Listing a large number of file is somewhat problematic and slow. Using the native C# library, running directly on an Azure VM, it can take several hours to list just a couple million files.
Switching from V1 to V2 requires the creation of a new Storage Account and that's pretty inconvenient.
Hadoop is organization-independent and can be used for various purposes ranging from archiving to reporting and can make use of economic, commodity hardware. There is also a lot of saving in terms of licensing costs - since most of the Hadoop ecosystem is available as open-source and is free
Great! Hadoop has an easy to use interface that mimics most other data warehouses. You can access your data via SQL and have it display in a terminal before exporting it to your business intelligence platform of choice. Of course, for smaller data sets, you can also export it to Microsoft Excel.
We went with a third party for support, i.e., consultant. Had we gone with Azure or Cloudera, we would have obtained support directly from the vendor. my rating is more on the third party we selected and doesn't reflect the overall support available for Hadoop. I think we could have done better in our selection process, however, we were trying to use an already approved vendor within our organization. There is plenty of self-help available for Hadoop online.
I feel that this is a highly reliable and scalable solution computing technology that is highly capable of processing large data sets across multiple servers and thousands of machines in a well-defined and distributed manner. Apache Hadoop can automatically scale up the number of servers and machines that are needed to process, store, and analyze data sets. It also handles explosions in data with big data technology. Apache Hadoop is good at handling all node failures as well.
The Azure Data Lake solution is designed for organizations that want to take advantage of big data. It provides a data platform that can help developers, data scientists, and analysts store data of any size and format and perform all types of processing and analytics across multiple platforms and programming languages. It can work with your existing solutions, such as identity management and security solutions. It also integrates with other data warehouses and cloud environments. It can be useful for organizations that need the above softwares.
As it was open source makes it popular choice for handling large chuck of datasets
It was free earlier but now it’s licensed but still enterprise is a fine tuned version which makes it easier for new users and administrators to use it
Our investment is worth every single penny.
Initial cost is more as you might need to hire administrators to setup the cluster and make them in scalable. But once done it’s pretty easy
The cost can be high for more advanced work. In some cases, for instance, time limits and lab runtimes may be too short if you are too slow to learn what is explained as you go along.
promote flexible team communication. You can create different spaces for different teams, and share files and tasks.