Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

30 August 2017

From Big Data Analytics to Big Tool Stack

Are we still discussing Big Data or we should start reviewing and monitoring its different tools and platforms? Also, how organization should determine the right tool to utilize as part of its data analytics architecture?

In today’s competitive business environment, Internet of Things devices and Information Services increasingly produce large amount of data in disparate structures. Many Open source and commercial tools continue to pop up to deal with the different characteristics of Big Data. As a result, there is an abundance of tools and platforms to analyze Big Data or act as building blocks of such. Just by reviewing open source tools, we have come across 300 tools. That number is final after we applied strict filters like legitimity of the source, license type, and last commitment activity. We are not talking about Big Data anymore, we are talking about Big Tools.

To extract value from Big Data, an organization should determine the right tool to utilize as part of its data analytics architecture. The right tool would depend on the characteristic of the data to be analyzed and the domain that the organization is operating under. The organization would train its IT workforce to obtain the technical expertise to be effective with those tools. Businesses incur costs when they try to adopt these tools or change their existing source codes to run on newer versions. In other words, technical debt. In the Big Tools era, there is no standard on how these tools come together and compose a data analytics architecture. Most of these tools are unknown to business world, some of these tools even we didn’t know. To illustrate this, Apache Beam and Apache SAMOA are good examples. Latest trends in the big data domain is moving towards providing a level of abstraction to utilize popular data processing platforms. Apache Beam implements its dataflow programming model on multiple processing platforms like Apache Spark and Apache Flink. Apache SAMOA enables programmers to apply machine learning algorithms on data streams. Applications developed with SAMOA can be executed on Apache Storm and Apache Samza. Moreover, new models and tools continue to emerge at a fast pace in Big Data domain. There is no established method to track the newest developments particularly for the open source tools.

We are working towards developing an open source big data analytics architecture. We are trying to keep it as simple as possible to provide a comprehensive picture on big data analytics lifecycle. For academia, the architecture will provide the state of the art, tools that are missing, and tools that are mature enough to be used as part of a research. It will also provide the method for tracking notable new open source tools popping up in different sources. For technical people, it will help determine the tool to use for a particular implementation. Small and medium sized enterprises can provide services using some these tools addressing the gaps in a bigger architecture. For an established firm trying to develop a strategy, the architecture will provide the comprehensive picture on what fits where. Commercial big data solution providers can also benefit from this architecture. They will see the capability they lack and collaborate with a small sized enterprise to provide that capability.

Mert Gokalpand Keres Kayabay are working with Mohamed Zaki to build this architecture. We will publish a working paper soon on this topic. 

23 January 2016

Enabling the 4th industrial revolution - "industrie 4.0" or the "internet of things"?

I've been struck recently by the range of people talking about new digital and data developments in manufacturing. Of particular interest has been the apparent explosion of discussion about industrie 4.0 (which is extremely popular in Germany), internet plus (which is being pushed by China) and the industrial internet (being promoted by GE among others).

Managers, consultants, policy makers and academics are all getting very excited about the potential of connected devices. The basic idea is that increasingly things (of all types) will be stuffed with sensors and connected to the internet. They will stream data back to the original equipment manufacturers who in turn will use sophisticated analytics to analyse and interpret the data. There are loads of examples. Caterpillar streams data back from mining and construction equipment, using this both to monitor the health of individual machines and also to identify ways in which productivity and efficiency might be increased. Rolls Royce monitors aero engines in flight, using sensors to track vibrations in fan blades, which allows them to predict whether or not maintenance is required. In the consumer world - wearable devices (e.g. Nike's fitbit or Garmin's forerunner) track and record exercise levels with the data being uploaded to the internet for benchmarking and comparison purposes.

One thing that I find interesting is the rate at which some of these ideas are developing and the level of interest there is in them. A good way of looking at this is to explore Google Trends, which basically tracks the popularity of search terms and plots these over time. Figure 1 shows a comparison of "industrie 4.0" and the "industrial internet". It neatly shows how effective the German Government and large industrial firms (including Bosch and Siemens) have been at promoting their vision of the future - industrie 4.0 - with a rapid rise of interest in industrie 4.0 since 2012.

 
Figure 1: Google Trends - Popularity of Search Terms "Industrie 4.0" and "Industrial Internet".

One could argue that industrie 4.0 is not a new vision. As Figure 1 also shows there has been interest in the industrial internet for at least a decade and indeed my colleagues at Cambridge IfM, most notably in DIAL (the Distributed Information and Automation Laboratory led by Professor Duncan McFarlane) have been getting our students to build demonstrators and simulations of intelligent factories for years. However, the recent excitement is a testament to the growing maturity of the technology and underlying data infrastructures that will enable a wider adoption of industrie 4.0 and this excitement has driven significant Government and policy interest, as well as research and development investment.

So is industrie 4.0 the answer? Are smart factories where materials and machines seamlessly collaborate to drive productivity and efficiency the future? I think the answer is "yes" and "no".  Much of the discussion about industrie 4.0 is still very internally focused - its a factory view of the world. A recent YouTube video illustrates the point. The video talks about a vision of tomorrow - the factory of the future - where machines and materials will use wireless data infrastructures to communicate and coordinate their activities. Yet the examples I started with are ones where the product has left the factory - manufacturers are worrying about how they can track their products once they go out into the field and are used in mines and quarries, on the wings of plans, or in our houses and cars. Here I would argue there is scope for a bigger and more impactful industrial revolution. The fourth industrial revolution will not just be about what happens inside factories, but it will encompass the entire value chain. It will involve remotely monitoring products as they are used in the field. Data will be collected and streamed back to original equipment manufacturers who will use these data to assess the health of assets, to determine whether any maintenance is required, to predict potential product breakdowns and failures. They'll use the data to improve the next generation of design, learning from experience. They'll use the data to look at how the customer's operation might be optimised. By gathering data from multiple machines in a quarry its possible to build a system model of the quarry and identify where bottlenecks lie and hence how productivity can be improved.

This extended view of the fourth industrial revolution won't just be enabled by industrie 4.0, but by the "internet of things" and that's why when you add "internet of things" to the Google Trends data a rather different picture emerges. Its clear that industrie 4.0 and the industrial internet are important component parts, but the real key to driving future success in manufacturing lies beyond the factory walls and this will be enabled by the internet of things.

 
Figure 2: Google Trends - Popularity of Search Terms Including "Internet of Things".

4 March 2015

Data-Driven Business Models (DDBM): A Blueprint for Innovation

“A Blueprint that can be utilized by established organisations to create their own business models that rely on data as key resource”

We live in a world where data is often described as the new oil. Just as with oil, the value contained within data is universally recognized. As the seemingly relentless march of big data into so many aspects of the commercial and non-commercial world continues, the practicalities of constructing and implementing data-driven business models (DDBMs) has become an ever-more important area of study and application. For today’s businesses, effective data utilization is concerned with not only competitiveness but also survival itself. In some industries, such as publishing, big data has spawned entirely new business models. For example, after a movement towards a digitally oriented distribution model and dwindling advertising revenues, certain publishers began to accumulate data relating to their online users – users whose demographic was particularly attractive to advertisers. This data could then be sold, enabling targeted and more effective advertising.

However, although big-data-oriented publications agree on the potentially positive impact of big data utilization, very few suggest how, in practice, it can be attained and none offer a research-based guide or blueprint that can be utilized by an existing business to help create and implement its own DDBM. The DDBM blueprint and the corresponding six fundamental questions of a data-driven business will allow existing businesses and start-ups to follow a step-by-step process to construct their own DDBM centred around the businesses’ own desired outcomes, organization dynamics, resources, skills and the business sector within which they sit. We are presenting an integrated framework that could help stimulate an organization to become data-driven by enabling it to construct its own DDBM in coordination with the six fundamental questions for a data-driven business.
  





The DDBM Blueprint suggests that creating a business model for a data-driven business involves answering six fundamental questions:

1. What do we want to achieve by using big data?
In order for a business to effectively utilize big data it is vital that its aims are clear and realistically attainable. Often an organization understands the potential value and benefit associated with data but fails to determine a specific aim before undertaking a time-consuming and costly data acquisition and analysis process. Seven key competitive advantages are attained; shortened supply chain, expansion, consolidation, processing speed, differentiation and brand. For example, the fashion retailer Zara aimed to achieve close to real-time customer insight into fashion industry trends and purchasing patterns so that it could better align itself with its customers, resulting in increased retail sales volume. Zara aimed to utilize a shortened supply chain to gain competitive advantage and incorporated near real-time sales statistics, blog posts and social media data into its analytic systems, to rush emerging trends to market.

2. What is our desired offering?
A business must decide in what way the DDBM construct will benefit the company’s current offering or, alternatively, create an entirely new one. Established businesses have a tendency to utilize data to improve or enhance their current customer offering, which is often called a ‘value proposition’.. A company can offer raw data that is primarily ‘a set of facts’ without an attached meaning. When data has been interpreted it becomes information or knowledge. Typically the output of any analytics activity attaches some insight or application.  For example, the mobile phone service provider AT&T increased the positive public perception of its brand after evaluating a customer sentiment analysis based upon both internal (current users) and external (potential users) data sources. This insight enabled AT&T to improve its product and service offering in areas considered most important to its potential and actual customers, thus maximizing the derived benefit from the investment.

3. What data do we require and how are we going to acquire it?
Data is obviously fundamental to a DDBM. Deciding which data is most applicable, and the nature of that data’s acquisition, is pivotally important to the success of a DDBM construction. Established businesses with a substantial number of customers, and therefore potential customer interaction points, are well positioned to effectively utilize customer-provided data within their DDBM, although this data is often combined with data from other sources. This high utilization of all available data sources by established organizations is indicative that these organizations understand the value of data and orient themselves towards becoming data-driven. For example, the fashion retailer Topshop combines customer-provided data, free available data from fashion blogs and social media, and existing data within its own databases when running predictive and descriptive analytics protocols to determine emerging trends within the highly competitive retail clothing industry

4. In what ways are we going to process and apply this data?
Methods of processing reveal the true value contained within data. Knowing which key activities will be utilized to process data enables the business to plan accordingly, ensuring that the necessary hardware, software and employee skill sets are in place. To develop a complete picture of the key activities, the different activities were structured along the steps of the ‘virtual value chain’. To gather data, a company can either generate the data itself internally or obtain the data from any external source (data acquisition). The generation can be done in various ways, either manually by internal staff, automatically through the use of sensors and tracking tools (e.g. Web-tracking scripts) or using crowd-sourcing tools. Insight is generated through analytics, which can be subdivided into: descriptive analytics, analytics activities that explain the past; predictive analytics, which predict/forecast future outcome; and prescriptive analytics, which predict future outcome and suggest decisions. In the financial services sector, where finely-tuned predictive analytic modelling influences business decisions, Goldman Sachs plans years in advance to ensure it has the capacity, hardware, processes and employee skill sets available to utilize increased data volumes and new technologies. In fact, approximately 30 per cent of all Goldman Sachs’ employees work in technology and development.

5. How are we going to monetize it?
Without the target of a quantifiable benefit to a business it is difficult to justify DDBM construction and implementation. Incorporating a revenue model into a DDBM is integral to its operational success. Seven revenue streams are identified: asset sale, giving away the ownership rights of a good or service in exchange for money; lending/renting/leasing, temporarily granting someone the exclusive right to use an asset for a defined period of time; licensing, granting permission to use a protected intellectual property like a patent or copyright in exchange for a licensing fee; a usage fee is charged for the use of a particular service; a subscription fee is charged for the use of the service; a brokerage fee is charged for an intermediate service; or advertising. Revenue models associated with a DDBM differ considerably from a standard subscription fee such as The New York Times for advertising. These models vary considerably between sectors and within industries.

6. What are the barriers to us accomplishing our goal?
Interestingly, our research and analysis revealed clear links between specific inhibitors to the implementation of a DDBM. Established businesses are experiencing cultural issues, personnel issues, and internal value perception obstacles to implementing a DDBM. Our study suggests that issues with personnel may be the most severe DDBM implementation inhibitors experienced by both new and established businesses and may be linked to a variety of other obstacles to a business becoming data-enabled.



1 March 2014

The Big Data Revolution: What Happened to Data Quality?

There's a wonderful irony in the world of Big Data Analytics. At a time when interest in Big Data appears to be growing exponentially, it appears that some are forgetting the fundamental challenges of Data Quality. A quick Google Trends analysis highlights the point. The chart below shows two trend lines extracted from Google Trends. The line in blue reflects the popularity of searches for Big Data, while the line in red shows the popularity of searches for Data Quality. It is important to note that the lines show relative popularity, not absolute volumes of search terms. In fact, Google Keywords suggests that in absolute terms searches for Big Data are about 20 times as popular as searches for Data Quality.


This raises an interesting question - what's happened to Data Quality? At a time when organisations are becoming ever more interested in using their data to create performance insights and predictions, interest the Data Quality appears to be declining. Is this because Data Quality is no longer an issue?

I don't think so. On three separate occasions in the last week alone I have been involved in discussions with senior managers from some of the world's leading manufacturing and service businesses. Each time, the issue of Data Quality has come up loud and clear. These firms recognise the potential of Big Data and Analytics, but are realistic enough to know that unless they sort out their data fundamentals - unless the track the right things and make sure the raw data if accessible and of high quality, all of the Big Data Analytics in the world is not going to help them. That's why - in the Cambridge Service Alliance - one of our projects this year is focusing on creating a data diagnostic - a methodology that can be used to check whether the data you have access to is appropriate and can be better used to optimise the delivery of your services and solutions. We're in the process of testing this data diagnostic at the moment and would love to hear from you if you'd be interested in being one of the pilot test sites.