Translate
Monday, December 7, 2020
Technology Initiatives
Tuesday, November 3, 2020
Interview Question - What are the 43 Vs in Big Data
If you are asked to explain the Vs in Data in an interview, what will be your answer. Is it 3Vs, 4Vs, 5Vs, 7Vs or 10 Vs? You might be surprised that you have 43 Vs!!!
Like the evolution of data, number of Vs have evolved over time.
Initially, it is 4Vs with Volume, Variety, Velocity and Veracity. Read detail at here.
Saturday, October 31, 2020
Do You Know - Big Data Facts
Not sure how accurate these facts are, but seems to be interesting.
1. Every two days we create as much information as we did from the beginning of time until 2003.
2. Over 90% of all the data in the world was created in the past 2 years.
3. By the end of the year 2020, digital information will be grown to 40 zettabytes. source
4. The total amount of data being captured and stored by industry doubles every 1.2 years. source
5. Every minute we send 204 million emails, generate 1.8 million Facebooks likes, send 278 thousand tweets, and upload 200,000 photos to Facebook. source
6. Google processes on average over 40 thousand search queries per second, making it over 1.5 billion in a single day. source
7. Around 100 hours of video are uploaded to Youtube every minute and it would take around 15 years to watch every video uploaded by users in one day. source
8. If you burned all of the data created in just one day on DVDs, you could stack them on top each other and reach the moon, Twice. source
9. AT&T is thought to hold the world largest volume of data in one unique database, its phone records database is 312 terabytes in size, and contains almost 2 trillions rows. source
10. 570 new websites spring into existence every minute of every day. source
11. Today's data centres occupy an area of land equal in size to almost 6,000 football fields. source
12. The NSA is thought to analyse 1.6% of all global internet traffic - around 30 petabytes every day. source
13. The Value of the Hadoop market is expected to soar from $2 billion in 2013 to $ 50 billion by 2020. source
14. The number of bits of information stored in the digital universe is thought to have exceeded the number of stars in the physical universe in 2007. source
15. The boom of the internet of things will mean that the amount of devices connected to the internet will rise to 50 billion by 2020. source
16. 12 million RFID tags used to capture data and track movement of objects in the physical world had been sold in by 2011. by 2021 it is estimated that number will have risen to 209 billion as the internet of things takes off. source
Friday, March 17, 2017
Microsoft Brings New Capabilities to HDInsight and DocumentDB
Wednesday, March 8, 2017
Confused with Big Data, Data Science and Data Analytics
This link will give you an idea about what are they, what skills you need and what are the salary levels.
Read more to get further details https://intellipaat.com/blog/compare-big-data-data-science-data-analytics/
Wednesday, June 29, 2016
Big Stories on Big Data
Read the full article at http://www.smartdatacollective.com/jessoaks11/330428/4-big-companies-using-big-data-successfully
Tuesday, September 29, 2015
Microsoft Readies Big Data Service That Includes New U-SQL Language
Earlier this year, Microsoft revealed plans to offer a new HDFS-compatible Hadoop File System data store that could run large analytics workloads called Azure Data Lake. So far, the technical preview hasn't appeared but the company today reiterated that the service, which it will actually call Azure Data Lake Store, will be available later this year and also announced some new services planned for its Azure-based Big Data portfolio.
Saturday, June 13, 2015
We Need Only Three Hadoops, And Maybe Three Systems
Enterprises like choices, they abhor vendor lock in, and they like the options that open source gives. But at the same time, too many choices fragments markets and doesn’t allow for the cultivation of larger vendors that can afford to take on the big problems. You need a balance between large enough to scale and so large as to be too overbearing, or sometimes, in the antitrust sense, to be abusive.
With the Hadoop Summit sponsored by Hortonworks going on this week and some tectonic changes in the systems market that have been going on for some months now, we are pondering how this all fits together given the current state of both Hadoop and the systems markets.
Read more http://www.theplatform.net/2015/06/11/we-need-only-three-hadoops-and-maybe-three-systems/
Wednesday, February 4, 2015
A Big Deal for Big Data? Microsoft, R and You
Microsoft announced at the end of last month that it was acquiring Revolution Analytics, a commercial provider of software and services for the open source R programming language. R is focused on statistical computing and predictive analytics, which will be key to enabling Microsoft to offer customers the kinds of tools they need to make sense of big data.
More at http://sqlmag.com/sql-server/big-deal-big-data-microsoft-r-and-you
Monday, February 2, 2015
Why big data matters to every business
Almost five years ago, Google’s Eric Schmidt announced that we had reached the point where more data was being created every two days, than in all of human history, up until 2003.
Read at http://www.hiscox.co.uk/business-blog/bernard-marr-column-big-data-matters-every-business/
Saturday, January 17, 2015
Big Data with the Microsoft Analytics Platform System
Want a broad-level technical look at Microsoft Analytics Platform System (APS), the high-performance and scalable solution built for modern data warehousing needs? If you're looking to improve your data loading and query response times (up to and beyond 100x over legacy solutions), get the details on this massively parallel processing appliance.
Experts highlight the hardware architecture and the software, and they walk you through an end-to-end demo of the appliance's power. They show you the Big Data capabilities in APS, with the included PolyBase, which can perform standard SQL queries to access and join Hadoop data with relational data. If you're working with real-time data analytics, you won't want to miss this course!
Friday, April 18, 2014
By the Numbers: The VAR - Big Data Market Opportunity
If you've been following the trends, you already know that Big Data can be very good for your business. But do you know just how good it can be? Let's take a look at some numbers, courtesy of the analysts.
Read more at here
Sunday, October 13, 2013
Ford drives in the right direction with big data
Nowadays, Ford uses big data to find out what their customers want and to develop better cars faster. Developing a product that requires 20.000 – 25.000 different parts to develop, big data seems to be the only way forward and Ford bets heavily on big data.
Ford actually opened a lab in Silicon Valley to improve its cars with big data. In order to progress their cars regarding fuel consumption, safety, quality and emissions, Ford gathers data from over four million cars with in-car sensors and remote application management software. All data is analysed in real-time giving engineers valuable information to notice and solve issues in real-time, know how the car responds in different road and weather conditions and any other forces that could affect the car.
Ford is also installing numerous sensors in their cars to monitor behaviour. They install over 74 sensors in cars including sonar, cameras, radar, accelerometers, temperature sensors and rain sensors. As a result, it Energi line of plug-in hybrid cars generate over 25 gigabytes of data every hour. This data is returned back to the factory for real-time analysis and returned to the driver via a mobile app. The cars in its testing facility even generate up to 250 gigabytes of data per hour from smart cameras and sensors.
Thursday, August 22, 2013
Five Myths About Big Data
Samuel Arbesman, an applied mathematician and network scientist, is a senior scholar at the Ewing Marion Kauffman Foundation and the author of “The Half-Life of Facts.” Follow him on Twitter: @Arbesman.
Big data holds the promise of harnessing huge amounts of information to help us better understand the world. But when talking about big data, there’s a tendency to fall into hyperbole. It is what compels contrarians to write such tweets as “Big Data, n.: the belief that any sufficiently large pile of s--- contains a pony.” Let’s deflate the hype.
1. “Big data” has a clear definition.
The term “big data” has been in circulation since at least the 1990s, when it is believed to have originated in Silicon Valley. IBM offers a seemingly simple definition: Big data is characterized by the four V’s of volume, variety, velocity and veracity. But the term is thrown around so often, in so many contexts — science, marketing, politics, sports — that its meaning has become vague and ambiguous.
2. Big data is new.
It’s true that today we can mine massive amounts of data — textual, social, scientific and otherwise — using complex algorithms and computer power. But big data has been around for a long time. It’s just that exhaustive datasets were more exhausting to compile and study in the days when “computer” meant a person who performed calculations.
Vast linguistic datasets, for example, go back nearly 800 years. Early biblical concordances — alphabetical indexes of words in the Bible, along with their context — allowed for some of the same types of analyses found in modern-day textual data-crunching.
The sciences also have been using big data for some time. In the early 1600s, Johannes Kepler used Tycho Brahe’s detailed astronomical dataset to elucidate certain laws of planetary motion. Astronomy in the age of the Sloan Digital Sky Survey is certainly different and more awesome, but it’s still astronomy.
3. Big data is revolutionary.
When a phenomenon or an effect is large, we usually don’t need huge amounts of data to recognize it (and science has traditionally focused on these large effects). As things become more subtle, bigger data helps. It can lead us to smaller pieces of knowledge: how to tailor a product or how to treat a disease a little bit better. If those bits can help lots of people, the effect may be large. But revolutionary for an individual? Probably not.
4. Bigger data is better.
In science, some admittedly mind-blowing big-data analyses are being done. In business, companies are being told to “embrace big data before your competitors do.” But big data is not automatically better.
Really big datasets can be a mess. Unless researchers and analysts can reduce the number of variables and make the data more manageable, they get quantity without a whole lot of quality. Give me some quality medium data over bad big data any day.
5. Big data means the end of scientific theories.
Chris Anderson argued in a 2008 Wired essay that big data renders the scientific method obsolete: Throw enough data at an advanced machine-learning technique, and all the correlations and relationships will simply jump out. We’ll understand everything.
But you can’t just go fishing for correlations and hope they will explain the world. If you’re not careful, you’ll end up with spurious correlations. Even more important, to contend with the “why” of things, we still need ideas, hypotheses and theories. If you don’t have good questions, your results can be silly and meaningless.
Sunday, May 26, 2013
How Vertica Was the Star of the Obama Campaign, and Other Revelations
The 2012 Obama re-election campaign has important implications for organizations that want to make better use of big data. The hype about its use of data is certainly justified, but a lesser-noticed aspect of the campaign ran against another kind of data hype we’ve all heard: the Silicon Valley hype around Hadoop that goes too far and claims an unreasonably large role for Hadoop. One of the most critical contributors to the Obama campaign’s success was the direct access it had to a massive database of voter data stored in Vertica.
Read more on http://citoresearch.com/data-science/how-vertica-was-star-obama-campaign-and-other-revelations
Sunday, March 31, 2013
Singapore NUS buys into Microsoft's big data vision
The National University of Singapore (NUS) today announced the successful implementation of SQL Server 2012 to give employees in its Centre for Instructional Technology (CIT) the ability to run hypotheses and validate IT proposals on a self-service basis.
The implementation of Microsoft's SQL Server 2012 took between three to six months, he revealed. Besides cutting down on time, the Power View dashboard for IVLE usage was significantly improved in that it was easily understood, configurable and managed almost entirely by end-users.
The Singapore university is one of the latest in the region to sign up for Microsoft's big data offerings. Arun Ulag, Microsoft's general manager for server & tools division in Asia-Pacific, said in the same briefing session Tuesday that Thailand's Department of Special Investigation (DSI) had implemented a big data system based on SQL Server 2012 and Apache Hadoop software to more effectively retrieve and correlate different sets of information stored up in its siloed databases.
Read more here
Friday, January 25, 2013
Hadoop + SQL Server + Excel = Big Data Analytics
Customers will need a modern data platform to evolve with the needs of the business and the data they are collecting. Big data has created a massive business opportunity for businesses around the world to find new, actionable insights from all the data they collect, whether structured or unstructured. Because at the end of the day, the biggest promise of Big Data is to drive smarter decisions from data. To do that you have to gain new insights from all types of data.
One of SQL Server 2012’s most significant differentiators to help enterprises handle large datasets from SQL Server 2008 is its compatibility with Hadoop. Hadoop allows users to process large amounts of both structured and unstructured data to quickly find insights on that data, and because Hadoop is open-source, it can provide these insights at a low cost.
Microsoft’s strategy seems to be to offer path of least resistance for it’s customers to adopting Big Data – by extending existing tools such as SQL Server and Office to work seamlessly with new data types and allowing companies to take advantage of their existing investments while making new ones.
Read more and related articles from here.




