Translate

Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Monday, December 7, 2020

Technology Initiatives

If someone asks what are the top three priorities what do you say? Is it Cloud, DevOps, Machine Learnings, IoT? The following is the survey done by Flexera for 303 respondents. 


Still, DevOps, Machine Learning, Big Data are not in the priorities list though many of us are taking on those topics. Digital transformation, Cybersecurity and Cloud migrations are in the top technology initiatives. 

  

Tuesday, November 3, 2020

Interview Question - What are the 43 Vs in Big Data

If you are asked to explain the Vs in Data in an interview, what will be your answer. Is it 3Vs, 4Vs, 5Vs, 7Vs or 10 Vs? You might be surprised that you have 43 Vs!!!

Like the evolution of data, number of Vs have evolved over time. 

Initially, it is 4Vs with Volume, Variety, Velocity and Veracity. Read detail at here

4Vs in Big Data
Source: https://www.ibmbigdatahub.com/infographic/extracting-business-value-4-vs-big-data

Then Value was added and it became 5Vs. Read details at here.
5Vs of Big Data
Source: https://www.edureka.co/blog/big-data-characteristics/

Then 5 become 8 with the additions Viscosity, Visualization, Virality, 

8Vs in Big Data
Source: https://twitter.com/vladobotsvadze/status/942197217473032192

It didn't take much longer to this to become 10 Vs. 

10Vs in Big Data
Source: http://houseofbots.com/news-detail/2819-1-the-10-v%27s-of-big-data

The latest article shows that we have 42 Vs which is not possible to draw in fancy diagrams. This article shows 42 different Vs as per in 2017. This article shows how the number of Vs have evolved over the years. 


Source: https://www.kdnuggets.com/2017/04/42-vs-big-data-data-science.html

During navigation through a Coursera course at  University of California San Diego, there is another V which is not in the 42. That is Valence. Valence in Big Data refers, to connectivity between data. When you are dealing with multiple data sets, there can be connectivity between data. When the Valence increases, you need to adapt to more complex algorithms. 

So we have 34 Vs of Properties of Big Data. Any things else you know. Surely, now this should have passed the hundred!

Saturday, October 31, 2020

Do You Know - Big Data Facts



Not sure how accurate these facts are, but seems to be interesting. 

1. Every two days we create as much information as we did from the beginning of time until 2003. 

2. Over 90% of all the data in the world was created in the past 2 years. 

3. By the end of the year 2020, digital information will be grown to 40 zettabytes. source

4. The total amount of data being captured and stored by industry doubles every 1.2 years. source

5. Every minute we send 204 million emails, generate 1.8 million Facebooks likes, send 278 thousand tweets, and upload 200,000 photos to Facebook. source

6. Google processes on average over 40 thousand search queries per second, making it over 1.5 billion in a single day. source 

7. Around 100 hours of video are uploaded to Youtube every minute and it would take around 15 years to watch every video uploaded by users in one day. source

8. If you burned all of the data created in just one day on DVDs, you could stack them on top each other and reach the moon, Twice. source 

9. AT&T is thought to hold the world largest volume of data in one unique database, its phone records database is 312 terabytes in size, and contains almost 2 trillions rows. source

10. 570 new websites spring into existence every minute of every day. source

11. Today's data centres occupy an area of land equal in size to almost 6,000 football fields. source 

12. The NSA is thought to analyse 1.6% of all global internet traffic - around 30 petabytes every day. source

13. The Value of the Hadoop market is expected to soar from $2 billion in 2013 to $ 50 billion by 2020. source

14. The number of bits of information stored in the digital universe is thought to have exceeded the number of stars in the physical universe in 2007. source

15. The boom of the internet of things will mean that the amount of devices connected to the internet will rise to 50 billion by 2020. source

16. 12 million RFID tags used to capture data and track movement of objects in the physical world had been sold in by 2011. by 2021 it is estimated that number will have risen to 209 billion as the internet of things takes off. source

Friday, March 17, 2017

Microsoft Brings New Capabilities to HDInsight and DocumentDB

Microsoft is headed to San Jose this week where they will be announcing new features for HDInsight and DocumentDB at Strata Hadoop + World. Additionally, the company is also announcing a new preview for SQL Server as well.
DocumentDB is Microsoft’s fully-managed NoSQL database service that is designed to help developers build highly scalable applications. The service, which offers guaranteed single-digit millisecond low latency at the 99th percentile, is adding support for Apache Spark.
Read more at https://www.petri.com/microsoft-brings-new-capabilities-hdinsight-documentdb
 

Wednesday, March 8, 2017

Confused with Big Data, Data Science and Data Analytics

Often people confused with Big Data, Data Science and Data Analytics. Are they same or different.
This link will give you an idea about what are they, what skills you need and what are the salary levels.

Read more to get further details https://intellipaat.com/blog/compare-big-data-data-science-data-analytics/

Wednesday, June 29, 2016

Big Stories on Big Data

Every one talks big on big data, But has it really implemented in real world. This article tells us about real world implementation on Big Data.

Read the full article at http://www.smartdatacollective.com/jessoaks11/330428/4-big-companies-using-big-data-successfully

Tuesday, September 29, 2015

Microsoft Readies Big Data Service That Includes New U-SQL Language

Earlier this year, Microsoft revealed plans to offer a new HDFS-compatible Hadoop File System data store that could run large analytics workloads called Azure Data Lake. So far, the technical preview hasn't appeared but the company today reiterated that the service, which it will actually call Azure Data Lake Store, will be available later this year and also announced some new services planned for its Azure-based Big Data portfolio.

More at https://redmondmag.com/blogs/the-schwartz-report/2015/09/microsoft-readies-new-big-data-services.aspx

Saturday, June 13, 2015

We Need Only Three Hadoops, And Maybe Three Systems

Enterprises like choices, they abhor vendor lock in, and they like the options that open source gives. But at the same time, too many choices fragments markets and doesn’t allow for the cultivation of larger vendors that can afford to take on the big problems. You need a balance between large enough to scale and so large as to be too overbearing, or sometimes, in the antitrust sense, to be abusive.

With the Hadoop Summit sponsored by Hortonworks going on this week and some tectonic changes in the systems market that have been going on for some months now, we are pondering how this all fits together given the current state of both Hadoop and the systems markets.

Read more http://www.theplatform.net/2015/06/11/we-need-only-three-hadoops-and-maybe-three-systems/

Wednesday, February 4, 2015

A Big Deal for Big Data? Microsoft, R and You

Microsoft announced at the end of last month that it was acquiring Revolution Analytics, a commercial provider of software and services for the open source R programming language. R is focused on statistical computing and predictive analytics, which will be key to enabling Microsoft to offer customers the kinds of tools they need to make sense of big data.

More at http://sqlmag.com/sql-server/big-deal-big-data-microsoft-r-and-you

Monday, February 2, 2015

Why big data matters to every business

Almost five years ago, Google’s Eric Schmidt announced that we had reached the point where more data was being created every two days, than in all of human history, up until 2003.

Read at http://www.hiscox.co.uk/business-blog/bernard-marr-column-big-data-matters-every-business/

Saturday, January 17, 2015

Big Data with the Microsoft Analytics Platform System

Want a broad-level technical look at Microsoft Analytics Platform System (APS), the high-performance and scalable solution built for modern data warehousing needs? If you're looking to improve your data loading and query response times (up to and beyond 100x over legacy solutions), get the details on this massively parallel processing appliance.
Experts highlight the hardware architecture and the software, and they walk you through an end-to-end demo of the appliance's power. They show you the Big Data capabilities in APS, with the included PolyBase, which can perform standard SQL queries to access and join Hadoop data with relational data. If you're working with real-time data analytics, you won't want to miss this course!

http://www.microsoftvirtualacademy.com/training-courses/big-data-with-the-microsoft-analytics-platform-system

Friday, April 18, 2014

By the Numbers: The VAR - Big Data Market Opportunity

If you've been following the trends, you already know that Big Data can be very good for your business. But do you know just how good it can be? Let's take a look at some numbers, courtesy of the analysts.

Read more at here

Sunday, October 13, 2013

Ford drives in the right direction with big data

Nowadays, Ford uses big data to find out what their customers want and to develop better cars faster. Developing a product that requires 20.000 – 25.000 different parts to develop, big data seems to be the only way forward and Ford bets heavily on big data.

Ford actually opened a lab in Silicon Valley to improve its cars with big data. In order to progress their cars regarding fuel consumption, safety, quality and emissions, Ford gathers data from over four million cars with in-car sensors and remote application management software. All data is analysed in real-time giving engineers valuable information to notice and solve issues in real-time, know how the car responds in different road and weather conditions and any other forces that could affect the car.

Ford is also installing numerous sensors in their cars to monitor behaviour. They install over 74 sensors in cars including sonar, cameras, radar, accelerometers, temperature sensors and rain sensors. As a result, it Energi line of plug-in hybrid cars generate over 25 gigabytes of data every hour. This data is returned back to the factory for real-time analysis and returned to the driver via a mobile app. The cars in its testing facility even generate up to 250 gigabytes of data per hour from smart cameras and sensors.

Read more at http://www.bigdata-startups.com/BigData-startup/ford-drives-direction-big-data/?goback=%2Egde_126490_member_5792682009092435971#%21

Thursday, August 22, 2013

Five Myths About Big Data

Samuel Arbesman, an applied mathematician and network scientist, is a senior scholar at the Ewing Marion Kauffman Foundation and the author of “The Half-Life of Facts.” Follow him on Twitter: @Arbesman.

Big data holds the promise of harnessing huge amounts of information to help us better understand the world. But when talking about big data, there’s a tendency to fall into hyperbole. It is what compels contrarians to write such tweets as “Big Data, n.: the belief that any sufficiently large pile of s--- contains a pony.” Let’s deflate the hype.

1. “Big data” has a clear definition.

The term “big data” has been in circulation since at least the 1990s, when it is believed to have originated in Silicon Valley. IBM offers a seemingly simple definition: Big data is characterized by the four V’s of volume, variety, velocity and veracity. But the term is thrown around so often, in so many contexts — science, marketing, politics, sports — that its meaning has become vague and ambiguous.

2. Big data is new.

It’s true that today we can mine massive amounts of data — textual, social, scientific and otherwise — using complex algorithms and computer power. But big data has been around for a long time. It’s just that exhaustive datasets were more exhausting to compile and study in the days when “computer” meant a person who performed calculations.

Vast linguistic datasets, for example, go back nearly 800 years. Early biblical concordances — alphabetical indexes of words in the Bible, along with their context — allowed for some of the same types of analyses found in modern-day textual data-crunching.

The sciences also have been using big data for some time. In the early 1600s, Johannes Kepler used Tycho Brahe’s detailed astronomical dataset to elucidate certain laws of planetary motion. Astronomy in the age of the Sloan Digital Sky Survey is certainly different and more awesome, but it’s still astronomy.

3. Big data is revolutionary.

When a phenomenon or an effect is large, we usually don’t need huge amounts of data to recognize it (and science has traditionally focused on these large effects). As things become more subtle, bigger data helps. It can lead us to smaller pieces of knowledge: how to tailor a product or how to treat a disease a little bit better. If those bits can help lots of people, the effect may be large. But revolutionary for an individual? Probably not.

4. Bigger data is better.

In science, some admittedly mind-blowing big-data analyses are being done. In business, companies are being told to “embrace big data before your competitors do.” But big data is not automatically better.

Really big datasets can be a mess. Unless researchers and analysts can reduce the number of variables and make the data more manageable, they get quantity without a whole lot of quality. Give me some quality medium data over bad big data any day.

5. Big data means the end of scientific theories.

Chris Anderson argued in a 2008 Wired essay that big data renders the scientific method obsolete: Throw enough data at an advanced machine-learning technique, and all the correlations and relationships will simply jump out. We’ll understand everything.

But you can’t just go fishing for correlations and hope they will explain the world. If you’re not careful, you’ll end up with spurious correlations. Even more important, to contend with the “why” of things, we still need ideas, hypotheses and theories. If you don’t have good questions, your results can be silly and meaningless.

Read more at http://www.washingtonpost.com/opinions/five-myths-about-big-data/2013/08/15/64a0dd0a-e044-11e2-963a-72d740e88c12_story.html

Sunday, May 26, 2013

How Vertica Was the Star of the Obama Campaign, and Other Revelations

The 2012 Obama re-election campaign has important implications for organizations that want to make better use of big data. The hype about its use of data is certainly justified, but a lesser-noticed aspect of the campaign ran against another kind of data hype we’ve all heard: the Silicon Valley hype around Hadoop that goes too far and claims an unreasonably large role for Hadoop. One of the most critical contributors to the Obama campaign’s success was the direct access it had to a massive database of voter data stored in Vertica.

Read more on http://citoresearch.com/data-science/how-vertica-was-star-obama-campaign-and-other-revelations

Sunday, March 31, 2013

Singapore NUS buys into Microsoft's big data vision

The National University of Singapore (NUS) today announced the successful implementation of SQL Server 2012 to give employees in its Centre for Instructional Technology (CIT) the ability to run hypotheses and validate IT proposals on a self-service basis.

The implementation of Microsoft's SQL Server 2012 took between three to six months, he revealed. Besides cutting down on time, the Power View dashboard for IVLE usage was significantly improved in that it was easily understood, configurable and managed almost entirely by end-users.

The Singapore university is one of the latest in the region to sign up for Microsoft's big data offerings. Arun Ulag, Microsoft's general manager for server & tools division in Asia-Pacific, said in the same briefing session Tuesday that Thailand's Department of Special Investigation (DSI) had implemented a big data system based on SQL Server 2012 and Apache Hadoop software to more effectively retrieve and correlate different sets of information stored up in its siloed databases.

Read more here

Friday, January 25, 2013

Hadoop + SQL Server + Excel = Big Data Analytics

Customers will need a modern data platform to evolve with the needs of the business and the data they are collecting. Big data has created a massive business opportunity for businesses around the world to find new, actionable insights from all the data they collect, whether structured or unstructured. Because at the end of the day, the biggest promise of Big Data is to drive smarter decisions from data. To do that you have to gain new insights from all types of data.

One of SQL Server 2012’s most significant differentiators to help enterprises handle large datasets from SQL Server 2008 is its compatibility with Hadoop. Hadoop allows users to process large amounts of both structured and unstructured data to quickly find insights on that data, and because Hadoop is open-source, it can provide these insights at a low cost.

Microsoft’s strategy seems to be to offer path of least resistance for it’s customers to adopting Big Data – by extending existing tools such as SQL Server and Office to work seamlessly with new data types and allowing companies to take advantage of their existing investments while making new ones.

Read more and related articles from here.