We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only.
Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
There are a few alternatives that can do the same transformation and aggregation like Apache Spark can do but most of them are not able to perform parallel computation. For example, pandas is a really good tool to do that but not parallelized; However, there are some tools that …
Apache Spark has much more better performance and features if we compare with Hive or map/reduce kind of solutions. Spark has many other features for machine learning, streaming.
1. Apache Spark is almost 100 % faster than Hadoop. 2. Apache Spark is more stable than Amazon EMR. 3. The end to end distributed machine library is more robust in Apache Spark.
Databricks uses Spark as a foundation, and is also a great platform. It does bring several add-ons, which we did not feel needed by the time we evaluated - and haven't needed since then. One interesting plus in our opinion was the engineering support, which is great depending …
It is easy to learn, read and to maintain. It brings the best of the Ruby on Rails framework from Java that helps to create a web service so easily. Communication is one of the most distinctive features of Apache Spark compared to alternative products. You are able to …
We evaluated SAS alongside with Apache Spark but during the course of proof of concept found that Apache Spark was able to support the hadoop eco-system and hadoop file system much better. It was much faster at that time while having the ability to process data quickly for the …
Consultor Tecnico - Java Developer and Php Developer.
Chose Apache Spark
I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
Even with Python, MapReduce is lengthy coding. Combination of Python with Apache Spark will not only shorten the code, but it will effectively increase the speed of algorithms. Occasionally, I use MapReduce, but Apache Spark will replace MapReduce very soon. It has many …
vs MapRedce, it was faster and easier to manage. Especially for Machine Learning, where MapReduce is lacking. Also Apache Storm was slower and didn't scale as much as Spark does. Spark elasticity was easier to apply compared to storm and MapReduce. managing resources for …
Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph) Python …
There are a few newer frameworks for general processing like Flink, Beam, frameworks for streaming like Samza and Storm, and traditional Map-Reduce. I think Spark is at a sweet spot where its clearly better than Map-Reduce for many workflows yet has gotten a good amount of …
Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
is much better as it’s easily accessible provides velvet documentation and fulfils all our needs as well as easily integrated into clients, environment
Google BigQuery is simpler and I say it has simpler UI too. If you have a clear long term ask , mainly business intelligence needs then Google BigQuery offers you good. If you need too much of features under a single cloud and you are ok to be lil clumsy then you can check …
I have used most of the data analytics platforms. Based on my work, I have found that the user interface of Google BigQuery is simple to navigate. I like the front view - ease of joining tables, and integration with other platforms.
Compared to every other analytics DB solution I've used, Google BigQuery was by far the easiest to set up and maintain, and scale. The price was also much lower for our use case (internal data analysis).
For our usage, Google BigQuery is cheaper and more performant. The others have their place, but in certain scenarios, Google BigQuery is a better solution.
We actually use Snowflake and BigQuery in tandem because they both currently meet various needs. Redshift, however, has barely been used since our migration away from it. In the case of both Snowflake and BigQuery, they beat Redshift by a long shot. The main reasons are their …
I came to use BigQuery from a traditional system like MS SQL server, the features which are available in BigQuery as a cloud service far outweigh the features from SQL server. I have not used other similar tools like Amazon Redshift but Google BigQuery serves multiple use cases …
Google BigQuery is cheaper and much faster as compared to both. While as compared to Snowflake , we tested it was faster and cheaper by 30%, that is after Snowflake tweaked their environment, if not for that it would have been 90% cheaper than snowflake. Redshift is not easy …
In my opinion, Google BigQuery is custom made to be the best data lake system that is easy to use, scalas to fit any business size, has inbuilt security, as well as tools for data integrity. Although a few other tools have some of the same functionality, Google BigQuery is the …
It's easier to connect data between BigQuery and looker studio instead of connecting the data between BigQuery and tableau in terms of data explore or dashboard creating. Therefore we are considering migrating dashboards from tableau to looker studio for the whole company. On …
When comparing Google BigQuery and Databricks, both platforms are powerful tools for managing and analyzing large datasets. BQ is ideal for businesses requiring large-scale analytics, reporting, and dashboarding with minimal operational overhead. It’s also great for ad-hoc …
Google BigQuery's main advantage over its direct competitors (Amazon Redshift and Azure Synapse) is that it is widely supported by non-Google software, while the others rely heavily on their own cloud ecosystems.
I have used other data manipulation tools like SQL Server and Google BigQuery feels more intuitive, Google provides so much documentation and tutorials that getting to know the software is not only easy but even satisfactory, so I'd say Google BigQuery is very superior to that …
Amazon Redshift was a likely alternative we were considering , but it needs to be provisioned on cluster and nodes, which increases infrastructure management, whereas Google BigQuery is serverless, so no infra management :) Also, I remember when comparing them we did found out …
Google BigQuery as a platform allows for more integrations and customizability than many other offerings. Users mostly need to understand the basics of database and SQL programming in order to get the most from the product. However, other products like Hevo do have less of a …
There are some areas in which this product is better while there are some in which others do better. It's not like Google BigQuery surpasses them in every metric. For a holistic view, I will say we use this because of - scalability, performance, ease of use, and seamless …
The data performance of Google BigQuery is best as per other software. Limitations on Google BigQuery's data size are superior to those of Microsoft SQL. Obtaining real-time data from several IoT devices is another benefit.
I personally find it by far simpler than Amazon redshift due it's onboarding seamlessness. For a quick start and simplify tye access to read the data big query provide better user experience and a smoother user interface. More importantly, the fact that Big Query can be easily …
BigQuery can automatically scale to accommodate the data and query load, providing potentially unlimited scalability. At the same time, Redshift requires manual scaling efforts to increase or decrease capacity, which might affect performance during scaling operations.
We focused more on data volume and less on full application capabilities. All in all, we found that the two solutions complement each other. For integration, some sources were better handled in SAP HANA, particularly other SAP systems where Google Big Query was more suitable …
SingleStore has a much lower query latency compared to BigQuery. Thus, we segregate faster tasks to SingleStore, and use BigQuery has our main database to store all historical data.
Google BigQuery i would say is better to use than AWS Redshift but not SQL products but this could be due to being more experience in Microsoft and AWS products. It would be really nice if it could use standard SQL server coding rather than having to learn another dialect of …
First and foremost, Google BigQuery's pricing structure, based on data processing and storage, is more cost-effective for our needs. Secondly, since we already use other Google Cloud services, its tight integration with them especially, with Cloud Storage and Dataflow was a big …
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
Google BigQuery is great for being the central datastore and entry point of data if you're on GCP. It seamlessly integrates with other Google products, meaning you can ingest data from other Google products with ease and little technical knowledge, and all of it is near real-time. Being serverless, BigQuery will scale with you, which means you don't have to worry about contention or spikes in demand/storage. This can, however, mean your costs can run away quickly or mount up at short notice.
It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Its serverless architecture and underlying Dremel technology are incredibly fast even on complex datasets. I can get answers to my questions almost instantly, without waiting hours for traditional data warehouses to churn through the data.
Previously, our data was scattered across various databases and spreadsheets and getting a holistic view was pretty difficult. Google BigQuery acts as a central repository and consolidates everything in one place to join data sets and find hidden patterns.
Running reports on our old systems used to take forever. Google BigQuery's crazy fast query speed lets us get insights from massive datasets in seconds.
It is challenging to predict costs due to BigQuery's pay-per-query pricing model. User-friendly cost estimation tools, along with improved budget alerting features, could help users better manage and predict expenses.
The BigQuery interface is less intuitive. A more user-friendly interface, enhanced documentation, and built-in tutorial systems could make BigQuery more accessible to a broader audience.
We have to use this product as its a 3rd party supplier choice to utilise this product for their data side backend so will not be likely we will move away from this product in the future unless the 3rd party supplier decides to change data vendors.
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
web UI is easy and convenient. Many RDBMS clients such as aqua data studio, Dbeaver data grid, and others connect. Range of well-documented APIs available. The range of features keeps expanding, increasing similar features to traditional RDBMS such as Oracle and DB2
I have never had any significant issues with Google Big Query. It always seems to be up and running properly when I need it. I cannot recall any times where I received any kind of application errors or unplanned outages. If there were any they were resolved quickly by my IT team so I didn't notice them.
I think Google Big Query's performance is in the acceptable range. Sometimes larger datasets are somewhat sluggish to load but for most of our applications it performs at a reasonable speed. We do have some reports that include a lot of complex calculations and others that run on granular store level data that so sometimes take a bit longer to load which can be frustrating.
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
BigQuery can be difficult to support because it is so solid as a product. Many of the issues you will see are related to your own data sets, however you may see issues importing data and managing jobs. If this occurs, it can be a challenge to get to speak to the correct person who can help you.
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Google BigQuery of course collects a much much larger array of raw data and can handle (practically) an unlimited amount of data. For a large enterprise like ours that relies on large-scale analytics, this is absolutely imperative. Google BigQuery can also combine GA4 data with external sources (like CRM tools), so our analytics can be unified. Due to our heavy reliance on GA4, Google BigQuery is the natural choice since it is a Google product and has better integration.
We have continued to expand out use of Google Big Query over the years. I'd say its flexibility and scalability is actually quite good. It also integrates well with other tools like Tableau and Power BI. It has served the needs of multiple data sources across multiple departments within my company.
Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.
In some places, Google BigQuery has helped us save some money by avoiding the need for expensive infrastructure and reducing some of the operational costs.
Scalability is up-to-date and really helpful in multiple places.
Knowledge transfer is easy as it is very user-friendly, so the learning curve has been reduced.
Also, it gives us more insights from our data, helping us make smarter decisions for our business.