Monday, 6 August 2018

"If you don't have a PhD, don't call yourself a data scientist"


“If you don’t have a PhD, don’t call yourself a data scientist”; with these remarks the government linked person set the stage to explain his views via a presentation on AI, and the problems implementation of AI suffers from today.

To flesh out his argument, he argued that only a PhD gives the rigour and access to large enough data to play with to become a data scientist.

A very interesting point of view from someone linked to the government.

Not everything he said was that controversial, at least to me.

The presenter used my favourite diagram for data science, the ‘Drew Conway’ diagram (1), acknowledging the importance of subject matter expertise. Data Science is a balanced combination of “Substantive Expertise” or subject matter expertise, “Maths and Statistics Knowledge” and “Hacking Skills” or IT skills.

Furthermore, the presenter also mentioned how hard it was to find all 3 skills at a required level in 1 person and also spoke of data science teams; or like what I say: “Data Science is a team Sport”.



Also the presenter was at pains to point out that a 3 month course in data science does not make you a data scientist, so even if you are an English Literature PhD, or hold a PhD in Astro Physics, a 3 months data science course does not make you a data scientist; it takes years.


I am on the wall on this one. I think “data science” like every subject needs practice, and while a 3 month course will most likely not give you enough experience, it doesn’t have to take years and years. Any expertise is gained through practice.

Furthermore, the presenter is a proponent of open source, and advises everyone to eschew classes and learn online instead, pay tens of dollars rather than hundreds. I am all for learning online, have taken classes from Data Camp (2) where I learnt a lot, as well as from Coursera (3).

But where it gets really weird, and please remember that the presenter is linked to the government, he then went on to “sell” is classroom courses, of around 3 months, and hopes he can provide some practical experience.

Unless he is targeting only PhDs as students, I find what he is saying quite contradictory...

The reason I mentioned English Literature and AstroPhysics is the presenter further mentioned that one of the reasons why the country may be finding it hard to find “data scientists” is the fault of HR departments. They are looking for a unicorn with degrees in computer science (let alone PhDs). The advice was that they should loosen the criteria and accept people from different disciplines and who have taken the online courses...

My view is not that dissimilar. I believe in passion and without knowing anything about a person, I would say that an engineer is more likely to make a good “data scientist” than a Statistician or a Computer Scientist. The reason is that to me, “data science” is about delivering value and the passion should be to solve problems, the end, not the means – AI/ML/Stats...

Then this goes back to the PhD question. Do I believe you can’t be a “data scientist” without a PhD? Well, it may be self-serving since my profile states “data scientist”, but no, I do not believe a PhD is required.

In fact, quite a few organisations have found this. Basically, people with PhDs are great at their own domain, but “data science: requires a multitude of skills that they may not have (for example subject matter expertise, or statistics for computer scientists, or IT skills for Statisticians) or may not want to engage in: the ‘dirty’ work of cleaning and preparing the data. Hence the organisations whose “data science” department is staffed purely by PhDs find it very difficult to get a decent RoI. (results, results and results).

While I am at it, I will also mention that another way that organisations get their staffing wrong (hey, may be that deserves a separate blog, but here goes) is in the fact that some “data scientists” delegate the data cleaning and preparation to “data preparation” or “data engineers”. It gets worse when the latter do not have a clear career path to the former, like sous-chefs becoming chefs... Data preparation should be done with a purpose, and unless the high and mighty “data scientist” can communicate the purpose effectively and in great detail (probably also requires some EQ), there is a risk that the data preparation will not be that fit for purpose.

Basically I believe that data cleaning and preparation is part of the role of a “data scientist” especially since “data science” is by nature iterative and iterations may involve obtaining and preparing data that was not included initially.

Quite a while ago I did an easy to understand view of the work of a unicorn (“data scientist”); as you can see, data preparation and transformation is part of the process. I can understand that someone who is good a solving business problems may not be very good at getting data in the most efficient way from various systems, or writing production ready code, but surely preparing data is part of the role after all, most people will tell you that this is 70%-80% of the work...(4)(5) 



So why I am upset enough to write this blog?

Basically I believe analytics/”Data Science” has the power to unlock enough value to create win-win (win) situations (organisation, customer/society, and (consultancy/ vendor)), and getting the framework for data science is critical in that regard.

From the presentation I attended, it would seem that the government has got some things right, some wrong, and some contradicting each other. I do hope they sort things out; unfortunately, the Peter principle may be at work.(6), or may be it’s HiPPOs (7) or both since there often is a high correlation between the two (A hippo named Peter...)

Actually ya, this might be the topic of my next blog, although I am also itching to write about AI/Automation and retraining...

P.S. Did I mention that the presenter said that SLR is part of AI?
(See it pays to read all the way to the end... now please clean the coffee from your device)

3.       https://www.coursera.org

Wednesday, 25 July 2018

Prudential Singapore (Insurance), Agents and Digitisation, some thoughts

A recent article (1) highlights Prudential’s  agents’ unhappiness with ‘digitisation’ of their business. The 2 main questions are whether this unhappiness is founded or not and what impact this has for the future of the insurance industry, especially in line with the emergence on Insurtech.

The complaints centre on the fact that Prudential is making available to customers via online channels the same products that agents offer to their prospects, including the basic/workhorse/most popular ones. Prudential management has been quick to state that the agents returns will not be affected.

So what is the most likely scenario?

Just think of it this way, if, without premiums going up, Pru is able to finance a digital platform, maintain it, and still pay agents full commission on sales they have nothing to do with, then it must mean that the premiums Pru customers are paying today are too high since they can pay for the extra costs of the platform...

Assuming that is not the case, so why would Pru be losing money? It is possible that Pru believes that agents do a terrible job of cross-selling products, and therefore the platform would pay for itself based in sales of products they wouldn’t otherwise have sold.

I worked in the insurance industry for a while, and of course there are customers who are more or less abandoned by agents and left to their own devices. Agents most often focus their efforts on a few customers whom they believe are more likely to buy or people they are closer to. But this doesn’t mean that simply having a platform where these less served customers could reach out to pru and make purchases would generate enough business (minus agents’ commissions) to fund the platform.

I would think that what agents are really worried about is the ‘customer ownership’. Again from my experience in the industry, agents believe it is them who own the customer relationship, not the insurance company. When an agent leaves, there is a process to deal with these ‘orphan’ accounts; they are basically reassigned to another agent, usually from the same agency.

It is possible that this may change, with the insurance company choosing to serve these customers via the platform rather than reassigning to agents, effectively having the insurance company now own the customer.

Furthermore, the company could enrol customers directly via the platform, removing agents from the picture for simple products, and using a pool of trained agents for products that require it.

In fact agents could fear that their role in the insurance industry will change drastically, and there may be a need for fewer of them, to act more as product specialists than engaging in the whole range of activities they cover today.

Does that mean that digitisation is not a good thing for agents?

The answer is no.

Digitisation is important, but it must be part of a transformation of the organisation towards being data driven. I believe that, in order for an organisation to become truly data driven, all its component parts have to be more or less  equally data driven – like a chain whose weakest link determined its strength, the least analytically mature component of an organisation determines its overall maturity and how far it is from being data driven.

I am an outsider, and I have no clue how analytically mature Pru is. However, as a customer of Pru, I get to see some of it. For example, I have received calls from Pru agents/agencies I have had no prior contact with, asking me about my policies with them; this should not happen. Usually only my agent or his agency should have access to what I purchased, not other agents/agencies. To me it means their databases are not that secure. The best part is that when I told my agent to complain to HQ, he basically said it happened all the time... So I am a bit skeptical about how data driven the organisation is.

Anyway, that being said, I think it would have made more sense for Pru to work together with their agents, rather than take the approach they have, which leads to some legitimate fears among their agency force.

Pru itself has said that Singaporeans are under insured (2), so if despite the personal attention that agents can give to their customers, these are still under-insured, what makes Pru think that the customer would by himself/herself buy from the platform?

Singapore is a relatively mature economy, insurers have been here for ages, and even then people are still under-insured. People do not decide to under-insure themselves, it’s most likely a matter of education, and that’s where the agent, if properly trained, can bridge the gap. People usually learn better in small settings; just like another Singaporean institution, the private tutor helps students in 1 on 1 or small class settings rather than in giant classrooms (or even worse on youtube videos...)

Basically, to me, agencies like Pru, still need their agency force to be on-board, they have a crucial role to play. (And my friendly agent would hit me with his hockey stick if I said otherwise). Education, and really doing proper financial planning for their customers is key. Sure there are stuff that can be bought ‘over the counter’ like travel insurance and are hugely successful (4). But not every insurance product is in that category.

Does that mean then that the platform is useless?

Of course not. I have even written how the whole insurance process could potentially be run on the cloud (3), so I get the importance of digitisation. To me the platform and analytics should support the agent and enable him/her to serve his/her customers better.

Properly built analytical models should give an idea of not only the potential need of a customer, but also the right timing. The simple way of working with agents to serve customers would be to:
  1. Inform the agents of whom, among their customers, is a good target for a specific offer now
  2. Allow the agent to personally contact some of the customers he/she wants to, get a commitment accordingly and a simple easy to use feedback loop.
  3. Contact the rest making reference to the agent if customers want to take action.

As a Pru customer, I have received SMSes that make me offers, sometimes they include my agent’s contact details in case I want to follow up. Well, none of their offers interested me; and the best part is that my agent knows that. We meet up once a year of so, and go over the policies, life… and he knows I do not need more coverage for now.

Is that all there is to digitisation or becoming data driven?

No. Far from it.

Insurance companies are made up of many parts, and the selling is only a small part of it. For example, once the customer decides to buy a policy (whether it is via a platform or an agent) how quickly is the policy sent back to the customer for take up?

Ideally this should all be straight through processing, especially for existing customers whose KYC (Know Your Customer) is still valid, all that’s needed is to get the customer to confirm the validity or make necessary updates. Then the insurer should have the validated data of teh customer and can proceed with the application proper.

Electronic forms are the best; while the platform should have this feature by default, it takes little to provide agents with the equipment necessary, a tablet for example (issues like online or offline are dependent on the specific market) with in-built checks to ensure all information necessary for an application is available and in correct form.

Then with correct use of technology, standard cases can be approved almost instantly; I have reproduced a diagram from my blog (3) below:



In such a case, underwriters need only work on exceptions. Using technology and analytics in this way makes the policy issuance process much more efficient, customers get their coverage faster and agents can focus on educating and selling rather than to have re-works, and the organisation gets policies in faster and or good enough quality. A win-win-win situation.

Similarly the claims process can be automated, again I would refer you to my earlier blog (3).

So the question is, has Pru done all this?

Well it has tried (5), but while this is a good beginning, this is far from customer centricity. Basically Pru gives discounts on premium if no claims are made. The interesting thing is that customers might choose the game the system, pay for small ailments rather than claim, enjoy lower premiums, and hit the insurer on the big ticket if any. I wonder whether behavioural changes have been taken into account in pricing: no claims doesn’t necessarily mean healthy, and incentivising gaming of the system is not usually a good idea. People are not stupid. Instead of piecemeal attempts, customers should be engaged, the organisation become customer centric rather than organisation or product centric as seems the case above.

And I didn’t even get into data driven customer centric  product design...

Sorry if this sounds like Pru bashing, it wasn;t my intention. But I see this episode as a case of an attempt at digitisation without looking at the big picture of being data driven and customer centric. I am sure Pru is not the only insurer in this situation. 

As long as there is no effort by insurers to become truly data driven and customer centric (and this is a process, not a single big bang), they will be vulnerable to more nimble technology driven players, picking off specific profitable segments, and that would be a double loss to the traditional insurers.

Friday, 13 July 2018

The world cup predictions prove: you need to use the right algo with the right data to have a chance at the right outcome.



I had written a simple blog about the English Premier League, using simple analytics to uncover some patterns and was planning to publish it. Then I thought I would see if the same patterns applied in the world cup. But have you seen so many worldcup predictions, and how spectacularly wrong they were? Hey, I can’t say anything about whether the octopus or the guinea pig can predict the winner, but I can certainly see when the tools/method/data used for prediction are unlikely to suit.

Let’s start with UBS(1)...

The technique used was simulation (Neymar’d be happy, so I guess Brazil would have an edge...). The first weird thing is that they decided to include Italy, a team that didn;t even qualify... Italy, again, a team that did not qualify, is ranked 12th...

This is where they got interesting,what did they simulate? Obviously not the tournament, since they have a phantom team. They input various valirables including the ELO ranking (2) of teams (claimed to be an objective measure of how good they are) into a statistical model and ran a series of monte carlo simulations.

The fun thing is that even while ELO ratings are ranking the performance of a team in a game, weighing the result by importance of the game (of course a win in a tournament is more important than in a friendly) among others. But interestingly, according to Wikipedia, “there is no single nor any official Elo ranking for football teams

It becomes hazier when these are claimed to be econometric methods (3).

Oh and by the way Italy was added to honour the nation...

Well if my aim was to come up with an accurate prediction I would try to make the best prediction possible, not use numbers that are not objective (although I may say they are), and include teams to honour them.


So, I would classify this prediction and a PR smoke and mirror exercise with some smoke coloured in teams colours to make you believe you are at a stadium. Don;t worry, there was very little violence at the World Cup, kudos to the Russian security.

Next, let’s go to Goldman Sachs...(4)

Now, no such old fashioned stuff like econometrics and monte carlo, too old and old school. Goldman Sachs went for everyone’s favourite: AI!

They used 200,000 models and simulate 1 million variations. Impressive. They used player and team attributes to forecast specific match scores, and then simulate the whole tournament. Now that makes more sense than having a phantom team...

Where is gets interesting is that they say “Brazil is expected to win its sixth World Cup title, defeating Germany in the final by an unrounded score of 1.70 to 1.41” mm... do they watch football? 1.71 to 1.41? is that 1-1 and goes to extra time or 2-1 and Brazil wins? Do you round up or down? Do they actually know there is extra time and penalty kicks?

Hmm... Looks like there is some domain knowledge lacking...




Fret not, Goldman Sachs was at it again (5). As the tournament progressed, they revised their predictions. After England beat Panama, they revised the prediction, now the final would be Brazil 1.59 England 1.17. hmmm still these pesky decimals... (also I guess it meant their initial model predicted England would not beat Panama...)

Still lacks understanding of the tournament, looks like all games are assumed to start with fresh squad... (at least equally fresh: no extra time)


But Goldman Sachs were not done yet! (6) After Belgium beat Brazil, they updated their prediction and made Belgium the favourite (32.6% chance of winning the world cup). Note that these predictions were made at semi-final stage; Goldman Sachs predicted the final would be Belgium England. Well they got it 100% wrong...






I would say Goldman has the right idea, simulate the whole tournament, use player and team stats... but not understanding football or at least the tournament game is not a good thing...

Any other method?

Well someone tried using graph theory (7).

I recommend this article, it is fun, easy to read and pretty.

But it eventually boils down to: the country whose players are playing in leagues where most players also play is most likely to win.  That’s what eigen vector centrality boils down to in this case. (ooops, sorry I try to avoid technical terms, but sometimes they come out)



This is a nice idea. If players at the world cup are the better ones, then the teams where many of them play would be of higher quality, Therefore, the teams with the more players in high quality teams are likely to have higher quality and therefore likely winners.


Makes sense, right? Especially if you take into account the fact that world cup teams are balanced (if you have 10 brilliant goal keepers in the toughest league, it won’t matter much here as every team can only have 3 goalkeepers). However it does ignore that football is a team sport, and although your squad matters very much in a tournament, you need some balance. If you have amazing talent in midfield but no strikers, who will score for you?

I would say I like this approach, can be made better with some football thinking; afterall football is a team sport; a team has to be more than the sum of individuals, the tactics matter, so does the manager who seems to have been ignored in all this.

So you  may ask, while I spit in everyone’s soup, is there a soup I would drink? The answer is yes! It may have sounded like a joke, but, to me, the most realistic prediction of the world cup that I have seen is a simulation of the whole tournament, with data on players, managers, tactics, from football manager (8)

Why?

It’s very simple, the data inside the game is of very high quality, the attributes of the players, their preferences, they style they play, how often they get injured... is all measured and included in the game; same for managers.

Furthermore, the whole tournament is simulated; has been 1000 times. And the results compiled.
And the likely winner is France.

To sum up, choosing the correct data to suit your problem (player and manager statistics, team statistics), the correct algorithms (simulating the whole tournament by the rules of the tournament) gives us a good chance to get the wanted outcome.

Hence France is likely to win the world cup, or so says the best analytical model I have seen.

But personally, I am likely to be rooting for Croatia, else my friend Genti may not forgive me ;)

And in case you are wondering, if you have seen my linkedin comments about the tournament (9), Tonton Zola Moukoko is a player where FM got it wrong; not everything is about data and sometimes life takes a turn (10)

  1. UBS predictions for the world cup https://www.businessinsider.sg/who-will-win-the-world-cup-2018-2018-5/?r=US&IR=T
  2. https://en.wikipedia.org/wiki/World_Football_Elo_Ratings
  3. https://www.cnbc.com/2018/05/17/world-cup-winner-predicted-by-ubs.html
  4. Goldman Sachs predictions for the world cup (https://www.businessinsider.sg/world-cup-predictions-pick-to-win-it-all-goldman-sachs-ai-model-2018-6/?r=UK&IR=T)
  5. Goldman Sachs predictions for the world cup  again https://www.ft.com/content/804a21be-7915-11e8-bc55-50daf11b720d
  6. Goldman Sachs predictions for the world cup  again and again https://www.businessinsider.sg/world-cup-predictions-goldman-sachs-ai-model-belgium-england-final/?r=US&IR=T
  7. https://cambridge-intelligence.com/graph-theory-world-cup-winner-prediction/
  8. https://www.youtube.com/watch?v=OxX_tdzFpgk
  9. https://www.linkedin.com/feed/update/urn:li:activity:6422819043379646464/
  10. https://offsiderulepodcast.com/2017/07/21/championship-manager-tonton-zola-moukoko/


Wednesday, 23 May 2018

Which gig/contract "data science" option is suitable for your organisation?


In my previous blog, I argued that few organisations should have full-time “data scientists” on their books. In this blog, inspired further by comments I received, I explore who the “data science” functions should be outsourced to.

To simplify things, I will start with 4 options.

Large Management Consultancies with “Data Science” arm
The first, most obvious, is to go with the big names; whether you look at the big management consultancies (Bain, BCG, McKinsey) to which you may add say Accenture or the big accountancy firms (Deloitte, EY and to a lesser degree PWC and KPMG) are all trying to establish a practice to gorge of the “data science” pie in the sky. So if you are from a huge organisation, then chances are, someone likely high up in your organisation has been approached by one or more of these.

I am definitely not against management consultancies getting into “data science”. In fact I believe that an organisation’s  “Analytical Maturity” is key to its ability to make use of data and become “data driven”. Therefore, the journey is much more than just applying some models/algorithms here and there, but the ability to consume and exploit them is critical. (In my previous blog (1), I kind of proxied that by “variety of ‘data science’ projects”).

The question is, can you afford these, is the RoI (Return on Investment) worth it...

Individual “data scientist”
On the other extreme, you can choose to pick an individual and either kick-start your analytical journey by showing quick-wins and good RoI, or start building an analytical culture from scratch – in which case the individual should be well rounded (it’s more akin to getting a sort of CDO).

Niche “data science” consultancies
A mid-way solution would be to engage niche consultancies specialised in “data science”. This is kind of best of both worlds (or worse); you get a group of people who may have complementary skills/specialisations, without the extreme overheads of layers of management and partners.

Technology Luminaries?
How about technology luminaries such as Cloudera, HortonWorks, MapR, Google, AliCloud... you may ask. Personally I think technology is a very important component of “data science” (afterall many algorithms have been available for years but compute capabilities required weren’t ready), however, technology is a tool. So, while I believe collaboration with these behemoths is a good way to go to equip your organisation, they should not be leading the efforts; carts and horses...

Positioning vis-a-vis Data Science Venn Diagram
If you look at the Drew Conway data science Venn Diagram (2), it becomes easier to see. The diagram below is inspired by the Drew Conway version, updated if you want.




Large Management or Accounting consultancies come from the area of Substantive Expertise, and in order to reach the data science have hired people with IT and Maths/Stats skills to form “Data Science” capable teams.

Individuals have specific skills, or a combination and can be coming from any area of the Venn Diagram (although, given the depth of knowledge and effort required to be in the centre of the diagram, they are less likely to be there)

Large Technology companies have an abundance of IT/Hacking Skills, and are picking up Maths/Stats, very often ignoring the “Substantive Expertise”, relying on “Machine Learning”/”AI”/”Deep Learning”. While it is debatable whether proponents/practitioners of machine learning today have enough maths/stats understanding, I think most would agree than there is a lack of business knowledge.

Some ML/AI/DL practitioners see this as a benefit, believing that all you need is data; I do not agree.

Small consultancies, like individuals can be anywhere in the diagram, but mostly are a combination of 2 aspects and most I am aware of are in the Machine Learning space, being staffed or started by people with strong Computer Science backgrounds. I am not saying that niche consultancies have no domain knowledge, they may have but either it comes from the IT side (as someone recently reminded me: working in the IT side of a bank doesn’t mean you understand how a bank works, but you understand how the IT in a bank works) or is quite specialised since the number of such experts is likely to be limited. They are after all niche.

What does this mean to your organisation?

If you are large enough and the kind of transformation you are willing to undertake is big enough, by all means engage a large management consultancy with a good “data science”/Analytics practice.

If your organisation is not that large, or prefers to spend more conservatively (could be the same amount but stretched over a longer period), then you could choose to find a niche consultancy and supplement them with your own substantive expertise (adding burden to your staff since the person should more or less be embedded with the niche vendor), the problem is that the talent pool is quite limited.

An alternative would be to constitute a team from individuals with various skill sets, you may even get an external expert with substantive expertise. The issue with this is that not everyone has the knowledge to hire specialists skills, and secondly screening, choosing people can be quite costly.
I am not sure why any organisation would look to technology companies to provide “data science”; it is just not their forte, and you would be better off pairing them with one of the above options to form a slightly more rounded team, more likely to reach the centre of the Venn Diagram.

Is there another option?

Well I expect Sesh (hi Sesh!) to say he knew it was coming, but frankly I didn’t . My blogs are usually written as I think about something, so they are quite raw (and the diagrams worse... The only blog that took me a lot of effort to write was the one about Hindu temple builders being data scientists). When I started thinking about this question, I genuinely expected to end up with conditions where each of the “data science” providers would have a place (except the large technology companies who really can only play an important supporting role, not a driving role). However ...

The long prescription

If your organisation is really large and serious about really transforming, then you should look at a large Management Consultancy with an analytics/”data science” arm.  They have the scale and capability to help you without adding much extra workload on your organisation, compared to cases when you would have to manage the process if you would engage niche consultancies or even individuals.

Organisation size is important – some of the large Management Consultancies would not even consider engaging with small organisations – but so is the analytical maturity of the organisation. Large Management Consultancies really come to the fore when there is a organisational transformation since they can leverage not only the analytics/”data science” arm, but also the traditional change management and associated skills where their traditional expertise lies.

If you are not large enough and/or you are not looking for organisational transformation, then you would be better off looking elsewhere. That is the whole point of analytics/”data science”; you run small experiments and keep what is good, chuck what is not. You do not need an army for that (as I mentioned in my earlier blog).

The advantage of working with individual “data scientists” is that you get dedicated people who you know or will get to know and who can fit well with your organisation (else you can always get someone else). The fit with the organisation can be experience in the specific subject area, ability to fit culturally with the organisation... and should not be under estimated.

However, no one individual will have the breadth of skills that you may need. Of course every data scientist can be adequate at a whole range of skills and subject areas, but no one can be a specialist at everything. Furthermore I believe analytics/”data science” is a team sport, you need more than just one person once you reach a certain level of maturity, especially when you are operationalising. 

Hence niche consultancies look good.

The main advantage of niche consultancies is that they have a decent breadth of skill-sets under one roof, and you may be able to access these skill-sets as and when you need them. Proper niche “data science” consultancies would at least a team of people covering the three circles of the Drew Conway diagram.

Looking at a “data science” project as a flow:


It takes a lot of different skill sets to have a successful “data science” project. For example solution architects to design data flows, database experts and data engineers to manage the data and especially for operationalisation, domain experts to interprete and craft the data, “data scientists” to build models/algorithms, visualisation experts to help take action among others... Not that these roles are a full list, nor that these are all specialised people, nor all are external to an organisation, but it just gives an idea of what is needed to operationalise “data science” and the advantage a niche consultancy would have over an individual, or a group of individuals brought together for a specific project.

However, when you get into consultancies, you enter the realm of overheads. While an individual has very little overheads, the larger the niche consultancy the more overheads: management, sales, adjustments for time on the bench, office space, administration expenses... Furthermore you may lose the fact that individual consultants, free from official partnerships/administrative work... are able to learn and grow in their areas of interest, as opposed to what the consultancy requires/prefers.

Furthermore, niche consultancies are niche because of some specialisation, be it by industry verticals or by function, or by a combination.  Therefore you need to find the right partner. One of the factors you may lose out is that one emerging trend is “data science” is cross-pollination where ideas and techniques from a different field/industry is modified and used.

So what is the solution?

The solution is to find networks of people who work together, of course preferably with an emphasis on analytics/”data science”. There are quite a few organisations that offer this “new” approach to analytics/”data science” and a few variations. My views are based on the ones I am familiar with.
Such a analytics/”data science” focused network basically combines the advantages of individual people with that of niche consultancies adding the potential for cross-pollination, without the disadvantages such as high overheads.

An organisation is free to choose the right person while that person benefits from being part of a network for learning and support, or a group of people with complementary skill sets to deliver an analytics/”data science” project without the huge overheads. Furthermore, all members have access to common resources and a support group of fellow experts just like larger formally organised consultancies.

What do the individual consultants get from joining such a network? Simple, the opportunity of working on projects they did not uncover themselves, of working on bigger projects than they could have by themselves, of learning from peers with similar mindset whether via discussions, training, or participating in projects in different roles.

Looks like I did end up with some simple rules of thumb...

In sum:
  1. If your organisation is large and serious about transformation, go for a large management consultancy with a “data science” practice, run a comprehensive transformation programme.
  2. If you have very specific needs and know an individual or a niche consultancy with reasonable overheads that exactly suits your these needs, then go for either of these options bearing in mind where you will need to manage/supplement.
  3.  For other cases, go for a network of analytics/”data science” experts that incorporates the advantages of the two options above without the disadvantages.
  4.  In general, use large technology vendors as providing technology rather than “data science” services, horses for courses.

So what does this mean for independent “data scientists” and niche consultancies?

Join an analytics/”data science” focused Network! A large portion of demand generation is via networks anyway, hence networking is not new to independent “data scientists” and niche consultancies. But it is an advantage to join more formal networks and get the benefits from there.
Individuals would benefit from joining networks by broadening their knowledge and gaining the ability of participating in larger projects. Note this does not have to mean loss of independence, in fact as long as there are no fees or other commitments; there is no downside for an individual to join a network.

Niche consultancies would still have high overheads, that’s because of their structure, but they would at least gain the ability to broaden the scope of projects they could take up by collaborating with other members of the network, allowing to continue specialisation while broadening the scope of projects that they can embark on.

Eventually, the choice of which network to join will be the critical one. The right network has to bring value to the individual or the niche consultancy. Value can be measured in many ways, and not all are purely monetary. As Doc argued (3), the values of the leader (and of the network) are critical.

P.S.
Personally I believe the labour market is changing so much that we will soon be back to the times where most of us would independently be selling our skills rather than being “full time employed” especially with benefits; back to middle ages/very early industrialisation.


Wednesday, 9 May 2018

Should you hire data scientists on gigs or as FTEs (Full Time Employees)?


Recently a conversation with a client caused me to re-think this situation. Yes, I fundamentally believe analytics/”Data science” is a comparative advantage an organisation can possess. However this need not mean an FTE (Full Time Employee) army, or does it?

As in most such questions, the answer is… "it depends".
To make things simple, I’ll consider 3 points of view:
A "Data Scientist"'s point of view
B Hiring Manager's point of view
C Organisation's point of view

Of course the ideal situation is where all 3 parties’ interests are in agreement.

I simplify the issue by measuring the status on 2 axes:
X.               Number of “data science” projects undertaken in a year
This is an indication of how often the skills of the “data scientist” are being used. This had to be projects that go beyond BI, need to have a predictive component. A “data scientist” is an expensive resource and it is a waste from all parties point of view to utilise a “data scientist” to build dashboards most of the time for example. There are other people much more skilled at this than “data scientists”.
Y.               Variety of projects undertaken in a year
The variety of projects is an indication of how far down the path of utilising “data science” or how broad the adoption or experimentation with “data science” is an organisation.

These 2 axes can be thought as components of analytical maturity – necessary but not sufficient conditions (I will discuss about analytical maturity in a subsequent blog). The more “data science’ projects you undertake in a year, the more likely you are to be using them. More importantly, the more varied they type of problems you are trying to solve using “data science”, the more likely it is that the adoption of “data science’ or the attempts at adoption permeate the organisation.

However, this looks at only the production side of things, not the consumption. There are many organisations out there who adopt a “if you build it…” and end up with white elephants.

At the end of this blog, I’ll describe an Occam’s Razor, but for now let’s look at things a bit more management school style.

A             “Data Scientist”'s Point of view


The top right is a sweet spot for the “Data Scientist”; he/she gets to continually learn things and use the skills in a variety of projects. This is a happily tired “Data Scientist” where it’s not work, just play.


The bottom right is where the “Data Scientist” is kept busy on projects that require his/her skills and knowledge, but these projects are repetitive. This can lead to boredom, and the palliative situation is to build strong feedback loops to keep improving and challenging the “Data Scientist”. Else it might make sense to rotate “data scientists’ and bring in new pairs of eyes to try and bring quantum improvements rather than continual polishing.

The top left is where the “Data scientist” gets a variety of projects but is under-utilised, a stop-start adventure. This is likely the case of an organisation who is starting with data science operationalisation, or is not mature enough to exploit “data science” fully. In this case, it might be better to hire specialist “data scientist” on a as-needed basis. This would be much better use of resources and make better use of “data scientists’” skills.

The bottom left is where no “data scientist” would want to tread; apparently he/she did not ask the right questions in the interview.

B             Hiring Manager's Point of View:

I was going to write an analysis of a hiring manager’s point of view, but I think it is enough to say that if they are not aligned to that of the organisation (principal agent problem) then it doesn’t matter what is really going on, all that matters is the impression you can give; it is an ego or resume padding trip and no reasoning can be applied.

C             Organisation’s point of view
The top-right is the sweet spot for the organisation; the “data scientist” is involved in many projects (well utilised) and a variety of projects (organisational analytical maturity), chances are organisations in this quadrant are able to generate sufficient RoI (Returns on Investment) from their “data scientist”. But does that mean that organisations in this quadrant should not hire data scientists on gigs? The answer depends what the “data scientist” can bring.

The bottom right is where the “data scientist” keeps doing the same things. Basically, at every iteration of a model, unless there have been structural changes, there will be incremental improvements. The question is whether this generates sufficient RoI for an FTE.

One of the questions clients (both when I was FTE and on gigs) often ask me is when they know a model needs to be relooked at. My usual answer is all to do with metrics. In the same way as I insist of proper performance metrics of models to be agreed at the start of a project (to ensure RoI), I also encourage clients to have acceptable variability in results. It’s a bit like having a point estimate and a band of acceptable intervals (CI). So only if the results degrade below an acceptable level would it be worth looking at (you cannot assume that the degradation is due to random fluctuations and the impact on returns is too negative).

So chances are, unless each application of the model generates huge returns, it might make more sense not to have a “data scientist” on the payroll but to use a “data scientist” on gigs (1).

The top left is where the “data scientist” is engaged on a variety of projects but not many projects. This is a clear case that a “data scientist” by gig is better. Not only do you utilise resources only when you need them, but you can ensure you get the best resource for the project.

The bottom left shows and organisation that has no use for a full time data scientist, but should explore “data science” via gigs.

Conclusion:

It is quite clear that the best case scenario for both the “data scientist” and the organisation  is when the “Data scientist” is fully engaged, building a variety of models/algorithms to solve different business use cases and generating RoI.

In the rest of the cases, the situation is unlikely to last. In cases where there is either variety or quantity but not both, the “data scientist” is likely to feel lack of growth and leave, causing the organisation to go through the expensive hiring process again and again, further impacting RoI.

In the case when there is neither quantity nor variety, there is no point engaging a full time “data scientist”. This is a very clear-cut case for having “data scientists” on a gig basis. The approach should be one of proof of concept, prove the RoI that can be obtained, using “data scientists” on gigs; not only is the cost overall lower, but the organisation can get specialists.

Simlarly, when there is variety but lack of quantity, “data scientists” on gigs offer the possibility of specialist help, and help only when needed. When you do not have projects requiring skills of a “data scientist” then why pay someone?

When there is no variety, the challenge is different. It is likely that this is the case when the analytical development of the organisation has stalled; for example, “data science” is used only in one area of the organisation, hence a lack of variety. Here again it might make more sense to only get “data scientists” on gig basis, to try and expand the variety and increase RoI by opening new avenues for returns. Also, another way of increasing RoI in this case if to hire “data scientists” for prototyping, but leaving maintenance to less expensive resources (2).

Finally, does that mean that if you have both variety and quantity you do not need “data scientists” on gigs? Well, it depends how well your RoI is doing. In the corporate world, where performance has to increase over time, using specialist help for prototyping, having a “different set of eyes” looking at business issues can be a solution. Of course you may choose to rotate your “data scientists” to ensure that freshness, but if could come at the cost of returns.

In a nutshell:

In a nutshell it all depends on the RoI. When I first took up a contract, it was very exciting as someone in the analytics field. My headcount was directly funded by the business and I had to justify my existence every year, I had specific RoI targets. That made me so alive. I guess that’s why I am believe strongly in RoI.

Apart of vanity and bragging rights on the part of management, why would an organisation spend on a resource that is generating low or no RoI?

Organisations have to measure the returns on their investments; that includes human capital. As long as the RoI is met, then hiring a FTE “data scientist” makes sense. Else, at the least, until the organisation matures enough to be able to allow the “data scientist” to generate that RoI, it should stick to “data scientists” on gigs.


1 There have been cases when clients want to put data scientists on a gig but also pay a retainer. This is sort of a compromise between the 2 models. However, retainers may not work. From the organisations perspective there will be an incentive to use the “data scientist” for non-“data science” work.
2 Another question I often get is “when do I need to review my models?”. I believe that the model metrics should not be decided just for a one time use, but also for on-going performance. For example, a targeted response rate of 15%, where review will take place if the response rate dips below say 12% for 2 consecutive runs.

Wednesday, 25 April 2018

Yes, facebook has taken liberties with the data they collect about you, but how safe is your DNA?




A while ago, I wrote about a new insurance product launched in Singapore that required you to submit your DNA as part of the deal – you got ‘personalised’ advice in exchange. The ad ridiculously showed two identical-looking twins receiving different advice (since identical twins share the same DNA...). (1) In that blog-post, I mentioned that the insurance company was at pains to stress that they had no access to the DNA, but I raised the prospect of someone buying that company collecting the DNA and not being bound by the same rules. And unfortunately this prospect is very real.

Let’s take a step back, am I talking about DNA or facebook?

These few weeks have been exciting for people interested in data and “Big Data”, since the extent of the data collected by Cambridge Analytica via facebook, very often without the subjects being aware (2). I have been going on about the need for us to own our data but this really takes the cake; you were not only giving away your data but that of your connections too (53 Australians took the test – and possibly gained something – but the data of 311,127 was harvested. Similarly 10 New Zealanders did so, and data from 63,724 as harvested. I am not saying there were national boundaries, but these numbers give an idea of the pandemic).

Ok, so people’s surfing habits, likes comments, photos they posted in public were accessed and used, but what use can be made of this data? As the time magazine article (1) mentioned, one use was for Mr Trump’s presidential campaign. And as this article shows, the efforts started in 2014 (4), and were very effective as confirmed by Mr Trump himself (5):
But they had this expression ‘drain the swamp.’ And I hated it, I thought it was so hokey. I said, ‘that is the hokiest, give me a break, I am embarrassed to say it.’ And I was in Florida where 25,000 people were going wild, and I said, ‘and we will drain the swamp’ — the place went crazy. I couldn’t believe it. And then the next speech I said it again and they went even crazier. ‘We will drain the swamp… we will drain the swamp,’ and every time I said it I got the biggest applause

So we can at least say that the data facebook ‘allowed’ Cambridge Analytica to harvest from the subjects was, at least, ‘useful’.

So what does that have to do with DNA?

Basically if you think that someone getting their hands on your surfing history and using it for their own purposes without your consent is bad, what if they get their hands on your DNA?

The organisation that holds the DNA for myDNA from Prudential is Prenetics Limited (7). Recently I read that Alibaba and Ping An insurance are the major investors in Prenetics (8). On one hand, I find it amusing that Ping An possibly have access to data that Prudential help collect. On the other I find it scary that the data of these people (of course I did not purchase myDNA) is now in the hands of another insurer.

Anyway, Prenetics claims that the DNA of over 200,000 people across South East Asia, China and Hong Kong were in their hands as early as October 2017 (9).

But, I am sure some nice people will say, there is a legitimate reason to do research into DNA; hospitals and universities have been doing so to the benefit of mankind for years. Yes, but I would argue that the CEOs of hospitals and universities have different experiences as compared to the CEO of Prenetics (Mr Danny Yeung) and that may affect how the data is being used:
Prenetics started out as ‘Multigene’ in 2009 when it span out from Hong Kong’s City University. Yeung joined the firm as CEO in 2014, after leaving Groupon following its acquisition of his Hong Kong startup uBuyiBuy, and it has been in startup mode since then. Prenetics has raised over $52 million from investors which, aside from Alibaba, include 500 Startups, Venturra Capital and Chinese insurance giant Ping An.”

This, I will admit, is pure speculation on my part. For all I know, Prenetics really wants to help mankind and bless everyone whose DNA they hold with better health and lower health care costs (prevention rather than cure). But I have other reasons to be sceptical.

Basically, even if humans ‘decoded’ the whole DNA sequence (which hasn’t been achieved yet (10)), even if you have inherited a predisposition to a condition, nobody can tell where you will actually get affected by it:
Genetic testing can provide only limited information about an inherited condition. The test often can't determine if a person will show symptoms of a disorder, how severe the symptoms will be, or whether the disorder will progress over time.” (11)

And to make things more interesting, the pieces of the genetic code that have not been sequences were considered useless or too hard to analyse given technological limitations, but are now being re-evaluated. Does that sound familiar? For people in the “Big Data” space (especially proponents of the “Data Lake”), it should.

One of the arguments of the “Data Lake” is that we do not know what data can be useful; even if we cannot extract is and use it now, we might as well keep it since it might be useful.
When I first started in this line of work, the kind of conversations I would have would be along these lines:
Q: “What data do you need?”
A: “Just give me what you have and I’ll analyse”
Q: “That is impossible, tell me what data do you need?”
A: “Ok, can I have the list of pieces of data that you have?”
Q: “That is impossible, tell me what you want and I will see if I have it...” ad nauseam

Now technology and acceptance of the usefulness of data have advanced and it is possible to “keep all the data” in a “Data Lake” or “Data Swamp” as some friends call it (Drain it! Drain it! Sorry I got caught for a moment).

Pieces of data that we would have had trouble analysing a few years ago such as weblogs, or pictures, or voice recordings can now be analysed relatively easily. But these pieces of data were routinely considered to be useless.

It is the same thing with DNA data. And to make it worse, there is the link between being at risk of some condition as per your DNA profile and actually getting that condition.

Basically, there is way too much data that would be needed to transform this ‘risk’ into something that can be measured with ‘enough accuracy’. That is what insurance companies try to do when they ask questions about your lifestyle, smoking, drinking... but these are very crude.

So is it fair that you could be penalised because of a feature of your DNA make-up? Are we slaves of our DNA?

What I am getting at is not the importance of DNA data, but rather at the care that must be taken when conclusions are made, and people penalised for things they may not be aware of.

To make things more fun, not only is Prenetics in China, Hong Kong and South East Asia, but it has recently acquired DNAFit (12). This impacts Prenetics in 2 ways. Firstly geographically, DNAFit’s market presence is mainly in Europe and is expanding to the USA. Secondly DNAFit goes direct to the consumer whereas Prenetics tended to reach the consumer via Insurance or Medical companies. (In fact even Linkedin is one of DNAFit’s customers).

The impact of direct-to-consumer DNAkits is debatable (13), but “a little learning is a dangerous thing” (14), add to this the emotional weight of ‘learning’ not necessary pleasant things about your own self...

So what I am saying is:
  1. As individuals we should have control over the data we produce by living (web/call/messaging behaviour, surveillance footage...
  2. But we should also have control over data we produce by existing (DNA).

I think there are many gaps between the general public (who have no issues with being facebook’s product in exchange for a quiz (15)) and those who have some idea of what can be done with such data; the same for DNA. And it is critical for people to be educated or educate themselves on this. As long as there is such an asymmetry of information, together with major issues with how people/machines use the data (people/machines, not technology or data itself), the cost of exploitation can be very high.

I would like to end this post with the poem by Alexander Pope (14):

A little learning is a dangerous thing ;
Drink deep, or taste not the Pierian spring :
There shallow draughts intoxicate the brain,
And drinking largely sobers us again.
Fired at first sight with what the Muse imparts,
In fearless youth we tempt the heights of Arts ;
While from the bounded level of our mind
Short views we take, nor see the lengths behind,
But, more advanced, behold with strange surprise
New distant scenes of endless science rise !
So pleased at first the towering Alps we try,
Mount o’er the vales, and seem to tread the sky ;
The eternal snows appear already past,
And the first clouds and mountains seem the last ;
But those attained, we tremble to survey
The growing labours of the lengthened way ;
The increasing prospect tires our wandering eyes,
Hills peep o’er hills, and Alps on Alps arise !


7 https://www.prudential.com.sg/en/prumydna/mydnapromotnc/ see point g: ““myDNA report” means the personalised report that Eligible Customers receive from Prenetics Limited”