TrustRadius: an HG Insights company

Apache Hadoop vs. Apache Spark

Save this comparison

Save this comparison

Add Product

Recommended Comparisons

    Overview
    ProductRatingMost Used ByProduct SummaryStarting Price

    Hadoop

    Score7.9 out of 10
    N/AHadoop is an open source software from Apache, supporting distributed processing and data storage. Hadoop is popular for its scalability, reliability, and functionality available across commoditized hardware.N/A

    Apache Spark

    Score9.2 out of 10
    N/AApache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.N/A
    Pricing
    HadoopApache Spark
    Editions & Modules
    No answers on this topic
    No answers on this topic
    Offerings
    Pricing Offerings
    HadoopApache Spark
    Free Trial
    NoNo
    Free/Freemium Version
    YesNo
    Premium Consulting/Integration Services
    NoNo
    Entry-level Setup FeeNo setup feeNo setup fee
    Additional Details——
    More Pricing Information
    Community Pulse
    HadoopApache Spark
    Considered Both Products
    Apache
    Chose Hadoop
    Apache Spark can be considered as an alternative because of its similar capabilities around processing and storing big data. The reason we went with Hadoop was the literature available online and integration capability with platforms like R Studio. The popularity of Hadoop has …
    Incentivized
    Chose Hadoop
    Apache Spark has an in memory processing model, making it powerful for lightning fast data processing. Apache Spark also exposes Scala and Python in APIs which is one of the most commonly used programming languages in data analytic and data processing domains.
    Incentivized
    Chose Hadoop
    Spark is a good alternative to Hadoop that can have faster querying and processing performance and can offer more flexibility in terms of applications that it can support.

    Google BigQuery has also been a great alternative and is especially great in terms of ease of use. The …
    Incentivized
    Chose Hadoop
    Hands down, Hadoop is less expensive than the other platforms we considered. Cloudera was easier to set up but the expense ruled it out. MS-SQL didn't have the performance we saw with the Hadoop clusters and was more expensive. We considered MS-SQL mainly for its ability …
    Incentivized
    Chose Hadoop
    • For real-time streaming, use Spark; can provide a stark contrast to the way MR works
    • Hadoop offers a scalable, cost-effective and highly available solution for big data storage and processing.
    • Amazon Redshift is somewhat closer to Hadoop. But to analyze Petabytes of data Hadoop …
    Incentivized
    Chose Hadoop
    • For real-time streaming, use Spark; can provide a stark contrast to the way MR works
    • Use Hive for querying purposes
    Incentivized
    Chose Hadoop
    Hadoop provides storage for large data sets and a powerful processing model to crunch and transform huge amounts of data. It does not assume the underlying hardware or infrastructure and enables the users to build data processing infrastructure from commodity hardware. All the …
    Incentivized
    Apache
    Chose Apache Spark
    Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
    Incentivized
    Chose Apache Spark
    • Apache Spark works in distributed mode using cluster
    • Informatica and Datastage cannot scale horizontally
    • We can write custom code in spark, whereas in Datastage and Informatica we can only choose the different features proivided already.
    Incentivized
    Chose Apache Spark
    Spark is simply awesome to work on with any data sets and also has an in-memory database which makes it very flexible.
    Incentivized
    Chose Apache Spark
    1. Apache Spark is almost 100 % faster than Hadoop.
    2. Apache Spark is more stable than Amazon EMR.
    3. The end to end distributed machine library is more robust in Apache Spark.
    Incentivized
    Chose Apache Spark
    I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
    Incentivized
    Chose Apache Spark
    All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
    Incentivized
    Chose Apache Spark
    Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
    Incentivized
    Chose Apache Spark
    Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph)
    Python …
    Incentivized
    Chose Apache Spark
    Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
    Incentivized
    Key User Insights
    Would buy again
    No answers on this topic
    No answers on this topic
    Delivers good value for the price
    No answers on this topic
    No answers on this topic
    Happy with the feature set
    No answers on this topic
    No answers on this topic
    Lived up to sales and marketing promises
    No answers on this topic
    No answers on this topic
    Implementation went as expected
    No answers on this topic
    No answers on this topic
    User Ratings
    HadoopApache Spark
    Likelihood to Recommend
    8.0
    (37 ratings)
    9.0
    (24 ratings)
    Likelihood to Renew
    9.6
    (8 ratings)
    10.0
    (1 ratings)
    Usability
    8.0
    (6 ratings)
    8.0
    (4 ratings)
    Performance
    8.0
    (1 ratings)
    -
    (0 ratings)
    Support Rating
    7.5
    (3 ratings)
    8.7
    (4 ratings)
    Online Training
    6.1
    (2 ratings)
    -
    (0 ratings)
    Data Sharing and Collaboration
    7.7
    (10 ratings)
    -
    (0 ratings)
    Data Sources
    8.7
    (10 ratings)
    -
    (0 ratings)