Wednesday, 13 May 2015



Unstructured Data and Data Hoarding: don't let the ogre scare you





 

This article in the local newspapers piqued my interest since it’s about data hoarding.

Step 1: "Discover users and resources - Determine what is important by rolling up your sleeves and digging through the piles of data" is probably the most critical step in the process. But who can tell what is useful, and especially what will be useful. 

There are 2 major contributors to volume of data: what to keep and for how long. (Embedded in ‘what to keep’ is also ‘in how much detail’ but let’s skip that for now)

For much analysis, a certain volume of data over time is required. You do not want to invest huge resources in a fad, especially if your time to market is not that rapid. Basically, to me, the length of time you need for analysis depends on how fast the environment you are playing in changes. Thus the length of time that we need to keep data is linked to the business you are in.

In terms of what to keep, there is much more debate. For example, how many people would have, a few years ago, decided that the stream of data that is constantly being output by sensors in a manufacturing process were worth keeping? They were consumed immediately and discarded. But nowadays, predictive maintenance is something that is relatively easy to do using precisely the huge volume of accurate sensor data that has been kept. And it is worth keeping this data since doing maintenance before a breakdown, even if it involves taking some components off-line, is much less costly that having to fix a broken machine and the associated impact on the production line.

Another example would be the huge volume of emails that employees engage in over time, especially changes in the patterns of these emails, in terms of frequency, direction, content, sentiment… The classic Enron email analysis is a clear example. How many organisations were analyzing employee email for more than flagging insider trading or breaches of corporate policy ‘now’, as opposed to understanding the patterns and detecting malpractice and collusion? Today this is a component of Human Resource Analytics.

I picked these 2 examples precisely because many people consider these (machine logs, body of emails) to be cases of unstructured data. While these may have been considered too difficult to use in the past, these are routinely used nowadays (hence some people prefer the term ‘semi-structured’). Today the frontier may lie at audio/video files, but I am sure everyone knows of cases where such data can be made very useful (and easily searchable in the case of audio files).

In sum, I would caution against underestimating Step 1. Discovering what is important is not a trivial task. The beauty of discovery is that we are only limited by our imaginations. Hence I would say: hoard as long as you think is relevant to your environment/market, and as much as possible, and don’t let the ogre of ‘unstructured’ scare you.

Tuesday, 21 April 2015

By all means experiment, but please do so intelligently


Experimentation and Design of Experiments

  

Experimentation is a ‘new’ buzzword, some even claiming “Big Data” demands experimentation.
I agree, and disagree. 

The need for experimentation has always been here – why would you launch a full scale campaign costing hundreds of thousands when you can spend some time refining and enhancing your offer and targeting?

It’s just that, with “Big Data”, data is available at a more granular level and hence we are able to detect more subtle events and sequences of events that can be seen as precursors to various courses of action. Today, organisations can quite easily know where their customers are, and allied with their recent behaviour, can create tailor-made offers. If I just had lunch at a fried chicken restaurant, and my bank has an offer at a yogurt bar nearby, it might be the exact thing to balance my tummy. Hence the bank could send me an offer for the yogurt bar.

Technology has helped in 2 further ways. Firstly, not only is “Big Data” available for analysis at a very granular level, but we are able to do so without having a Computer Science Degree. Hence the number of potential experiments that can be attempted has increased. Secondly, technology has also increased the ability to execute these experiments; apps, text messages are all avenues where we can be reached.

But what, to me, is critical is the design of the experiments. Just like with “Big Data” analytics, while technology has put the ability to analyse data in the hands of many, there are fundamentals that put probability on your side. Proper experimental design is critical and the costs of not doing so can be catastrophic.

The articles linked below highlight some failures in the medical field that have huge implications and that could be avoided simply with proper experimental design. The sad part is that the failures are so basic that a little thought would have avoided them:
                Ensure large enough a sample to conduct experiments (power tests)
                Do not bias the experiment by selecting groups to affect success probabilities 
(random selection)
                Do not bias the experiment by your own behaviour (double blind tests)

By all means experiment, but please do so intelligently.


Monday, 16 March 2015

How governance should take advantage of technology/"Big Data" and stop making excuses.


How governance should take advantage of technology/"Big Data" and stop making excuses.
( Why the 30000foot defense shouldn't be allowed in this world of "Big Data")

“Big Data” has been touted as the panacea, personally I think it is important to understand what is worth doing and what is not, to discuss what we should allow to be done, and what not, and be allowed to change our minds as time goes by as we see the consequences of our decisions.
One topic where, I believe, “Big Data” is very useful is in that of governance. 

The amount of data that is captured within an organization, together with the technologies that democratise the ability to extract or pursue meaning in a mass of data should make it impossible for a governance body to claim the 30000 foot defense: “we are too high up to have known”. No, sorry; you should have known (especially when you are paid very handsomely to simply ensure governance).

What can big data and appropriate visualisations allow you to see? Everything down to the minute details should you choose to:

With text mining you can scan through megabytes of documents and instantly get the gist of it all, allowing you to query further (yes, you can have ‘experts’ to do the querying if you want, but you cannot say you had no idea). For example the summarization of the Senate report on torture:
All you have to do when you see something you don’t find kosher is to ask: “tell me more about this”, and you drill down further.
Or if you want to understand how people’s or organisations’ behaviour varied, you could visually see how they differ (the example refers to an earlier piece on whiskies rather than people or organisations, but the idea is the same, comparing attributes across individuals):



And with some mouse work:
Ask “how does Mr. GlenFiddich differ from the rest?”.
Or what do the relationships between my clients/relationship managers look like?

And what is this interesting relationship in the centre?




In sum, with “Big Data”, there’s very little going on, especially within an organization, that can be hidden from people whose mandate is to ensure everything goes according to expected standards and values; unusual behaviour leaves traces that are so visible; you don’t need to be a rocket scientist to uncover them.

Wednesday, 4 February 2015

Visual Text Mining: the starfish of torture

There is so much information in text that it is sometimes difficult to choose what to read.

Text mining is a great way of extracting the main ideas from a body of text, but I think that it has to be allied with good visualisation, to capture and make it easy for anyone to understand the main ideas of a body of text.

The US Senate released a report on torture a few months ago, and I tried to gather the main conclusions and express them visually.

The summary is as follows, the starfish of torture:


I have created a deck of slides that illustrates the power of visualisation in bringing text mining to life and capturing ideas.