Wednesday, 23 May 2018

Which gig/contract "data science" option is suitable for your organisation?


In my previous blog, I argued that few organisations should have full-time “data scientists” on their books. In this blog, inspired further by comments I received, I explore who the “data science” functions should be outsourced to.

To simplify things, I will start with 4 options.

Large Management Consultancies with “Data Science” arm
The first, most obvious, is to go with the big names; whether you look at the big management consultancies (Bain, BCG, McKinsey) to which you may add say Accenture or the big accountancy firms (Deloitte, EY and to a lesser degree PWC and KPMG) are all trying to establish a practice to gorge of the “data science” pie in the sky. So if you are from a huge organisation, then chances are, someone likely high up in your organisation has been approached by one or more of these.

I am definitely not against management consultancies getting into “data science”. In fact I believe that an organisation’s  “Analytical Maturity” is key to its ability to make use of data and become “data driven”. Therefore, the journey is much more than just applying some models/algorithms here and there, but the ability to consume and exploit them is critical. (In my previous blog (1), I kind of proxied that by “variety of ‘data science’ projects”).

The question is, can you afford these, is the RoI (Return on Investment) worth it...

Individual “data scientist”
On the other extreme, you can choose to pick an individual and either kick-start your analytical journey by showing quick-wins and good RoI, or start building an analytical culture from scratch – in which case the individual should be well rounded (it’s more akin to getting a sort of CDO).

Niche “data science” consultancies
A mid-way solution would be to engage niche consultancies specialised in “data science”. This is kind of best of both worlds (or worse); you get a group of people who may have complementary skills/specialisations, without the extreme overheads of layers of management and partners.

Technology Luminaries?
How about technology luminaries such as Cloudera, HortonWorks, MapR, Google, AliCloud... you may ask. Personally I think technology is a very important component of “data science” (afterall many algorithms have been available for years but compute capabilities required weren’t ready), however, technology is a tool. So, while I believe collaboration with these behemoths is a good way to go to equip your organisation, they should not be leading the efforts; carts and horses...

Positioning vis-a-vis Data Science Venn Diagram
If you look at the Drew Conway data science Venn Diagram (2), it becomes easier to see. The diagram below is inspired by the Drew Conway version, updated if you want.




Large Management or Accounting consultancies come from the area of Substantive Expertise, and in order to reach the data science have hired people with IT and Maths/Stats skills to form “Data Science” capable teams.

Individuals have specific skills, or a combination and can be coming from any area of the Venn Diagram (although, given the depth of knowledge and effort required to be in the centre of the diagram, they are less likely to be there)

Large Technology companies have an abundance of IT/Hacking Skills, and are picking up Maths/Stats, very often ignoring the “Substantive Expertise”, relying on “Machine Learning”/”AI”/”Deep Learning”. While it is debatable whether proponents/practitioners of machine learning today have enough maths/stats understanding, I think most would agree than there is a lack of business knowledge.

Some ML/AI/DL practitioners see this as a benefit, believing that all you need is data; I do not agree.

Small consultancies, like individuals can be anywhere in the diagram, but mostly are a combination of 2 aspects and most I am aware of are in the Machine Learning space, being staffed or started by people with strong Computer Science backgrounds. I am not saying that niche consultancies have no domain knowledge, they may have but either it comes from the IT side (as someone recently reminded me: working in the IT side of a bank doesn’t mean you understand how a bank works, but you understand how the IT in a bank works) or is quite specialised since the number of such experts is likely to be limited. They are after all niche.

What does this mean to your organisation?

If you are large enough and the kind of transformation you are willing to undertake is big enough, by all means engage a large management consultancy with a good “data science”/Analytics practice.

If your organisation is not that large, or prefers to spend more conservatively (could be the same amount but stretched over a longer period), then you could choose to find a niche consultancy and supplement them with your own substantive expertise (adding burden to your staff since the person should more or less be embedded with the niche vendor), the problem is that the talent pool is quite limited.

An alternative would be to constitute a team from individuals with various skill sets, you may even get an external expert with substantive expertise. The issue with this is that not everyone has the knowledge to hire specialists skills, and secondly screening, choosing people can be quite costly.
I am not sure why any organisation would look to technology companies to provide “data science”; it is just not their forte, and you would be better off pairing them with one of the above options to form a slightly more rounded team, more likely to reach the centre of the Venn Diagram.

Is there another option?

Well I expect Sesh (hi Sesh!) to say he knew it was coming, but frankly I didn’t . My blogs are usually written as I think about something, so they are quite raw (and the diagrams worse... The only blog that took me a lot of effort to write was the one about Hindu temple builders being data scientists). When I started thinking about this question, I genuinely expected to end up with conditions where each of the “data science” providers would have a place (except the large technology companies who really can only play an important supporting role, not a driving role). However ...

The long prescription

If your organisation is really large and serious about really transforming, then you should look at a large Management Consultancy with an analytics/”data science” arm.  They have the scale and capability to help you without adding much extra workload on your organisation, compared to cases when you would have to manage the process if you would engage niche consultancies or even individuals.

Organisation size is important – some of the large Management Consultancies would not even consider engaging with small organisations – but so is the analytical maturity of the organisation. Large Management Consultancies really come to the fore when there is a organisational transformation since they can leverage not only the analytics/”data science” arm, but also the traditional change management and associated skills where their traditional expertise lies.

If you are not large enough and/or you are not looking for organisational transformation, then you would be better off looking elsewhere. That is the whole point of analytics/”data science”; you run small experiments and keep what is good, chuck what is not. You do not need an army for that (as I mentioned in my earlier blog).

The advantage of working with individual “data scientists” is that you get dedicated people who you know or will get to know and who can fit well with your organisation (else you can always get someone else). The fit with the organisation can be experience in the specific subject area, ability to fit culturally with the organisation... and should not be under estimated.

However, no one individual will have the breadth of skills that you may need. Of course every data scientist can be adequate at a whole range of skills and subject areas, but no one can be a specialist at everything. Furthermore I believe analytics/”data science” is a team sport, you need more than just one person once you reach a certain level of maturity, especially when you are operationalising. 

Hence niche consultancies look good.

The main advantage of niche consultancies is that they have a decent breadth of skill-sets under one roof, and you may be able to access these skill-sets as and when you need them. Proper niche “data science” consultancies would at least a team of people covering the three circles of the Drew Conway diagram.

Looking at a “data science” project as a flow:


It takes a lot of different skill sets to have a successful “data science” project. For example solution architects to design data flows, database experts and data engineers to manage the data and especially for operationalisation, domain experts to interprete and craft the data, “data scientists” to build models/algorithms, visualisation experts to help take action among others... Not that these roles are a full list, nor that these are all specialised people, nor all are external to an organisation, but it just gives an idea of what is needed to operationalise “data science” and the advantage a niche consultancy would have over an individual, or a group of individuals brought together for a specific project.

However, when you get into consultancies, you enter the realm of overheads. While an individual has very little overheads, the larger the niche consultancy the more overheads: management, sales, adjustments for time on the bench, office space, administration expenses... Furthermore you may lose the fact that individual consultants, free from official partnerships/administrative work... are able to learn and grow in their areas of interest, as opposed to what the consultancy requires/prefers.

Furthermore, niche consultancies are niche because of some specialisation, be it by industry verticals or by function, or by a combination.  Therefore you need to find the right partner. One of the factors you may lose out is that one emerging trend is “data science” is cross-pollination where ideas and techniques from a different field/industry is modified and used.

So what is the solution?

The solution is to find networks of people who work together, of course preferably with an emphasis on analytics/”data science”. There are quite a few organisations that offer this “new” approach to analytics/”data science” and a few variations. My views are based on the ones I am familiar with.
Such a analytics/”data science” focused network basically combines the advantages of individual people with that of niche consultancies adding the potential for cross-pollination, without the disadvantages such as high overheads.

An organisation is free to choose the right person while that person benefits from being part of a network for learning and support, or a group of people with complementary skill sets to deliver an analytics/”data science” project without the huge overheads. Furthermore, all members have access to common resources and a support group of fellow experts just like larger formally organised consultancies.

What do the individual consultants get from joining such a network? Simple, the opportunity of working on projects they did not uncover themselves, of working on bigger projects than they could have by themselves, of learning from peers with similar mindset whether via discussions, training, or participating in projects in different roles.

Looks like I did end up with some simple rules of thumb...

In sum:
  1. If your organisation is large and serious about transformation, go for a large management consultancy with a “data science” practice, run a comprehensive transformation programme.
  2. If you have very specific needs and know an individual or a niche consultancy with reasonable overheads that exactly suits your these needs, then go for either of these options bearing in mind where you will need to manage/supplement.
  3.  For other cases, go for a network of analytics/”data science” experts that incorporates the advantages of the two options above without the disadvantages.
  4.  In general, use large technology vendors as providing technology rather than “data science” services, horses for courses.

So what does this mean for independent “data scientists” and niche consultancies?

Join an analytics/”data science” focused Network! A large portion of demand generation is via networks anyway, hence networking is not new to independent “data scientists” and niche consultancies. But it is an advantage to join more formal networks and get the benefits from there.
Individuals would benefit from joining networks by broadening their knowledge and gaining the ability of participating in larger projects. Note this does not have to mean loss of independence, in fact as long as there are no fees or other commitments; there is no downside for an individual to join a network.

Niche consultancies would still have high overheads, that’s because of their structure, but they would at least gain the ability to broaden the scope of projects they could take up by collaborating with other members of the network, allowing to continue specialisation while broadening the scope of projects that they can embark on.

Eventually, the choice of which network to join will be the critical one. The right network has to bring value to the individual or the niche consultancy. Value can be measured in many ways, and not all are purely monetary. As Doc argued (3), the values of the leader (and of the network) are critical.

P.S.
Personally I believe the labour market is changing so much that we will soon be back to the times where most of us would independently be selling our skills rather than being “full time employed” especially with benefits; back to middle ages/very early industrialisation.


Wednesday, 9 May 2018

Should you hire data scientists on gigs or as FTEs (Full Time Employees)?


Recently a conversation with a client caused me to re-think this situation. Yes, I fundamentally believe analytics/”Data science” is a comparative advantage an organisation can possess. However this need not mean an FTE (Full Time Employee) army, or does it?

As in most such questions, the answer is… "it depends".
To make things simple, I’ll consider 3 points of view:
A "Data Scientist"'s point of view
B Hiring Manager's point of view
C Organisation's point of view

Of course the ideal situation is where all 3 parties’ interests are in agreement.

I simplify the issue by measuring the status on 2 axes:
X.               Number of “data science” projects undertaken in a year
This is an indication of how often the skills of the “data scientist” are being used. This had to be projects that go beyond BI, need to have a predictive component. A “data scientist” is an expensive resource and it is a waste from all parties point of view to utilise a “data scientist” to build dashboards most of the time for example. There are other people much more skilled at this than “data scientists”.
Y.               Variety of projects undertaken in a year
The variety of projects is an indication of how far down the path of utilising “data science” or how broad the adoption or experimentation with “data science” is an organisation.

These 2 axes can be thought as components of analytical maturity – necessary but not sufficient conditions (I will discuss about analytical maturity in a subsequent blog). The more “data science’ projects you undertake in a year, the more likely you are to be using them. More importantly, the more varied they type of problems you are trying to solve using “data science”, the more likely it is that the adoption of “data science’ or the attempts at adoption permeate the organisation.

However, this looks at only the production side of things, not the consumption. There are many organisations out there who adopt a “if you build it…” and end up with white elephants.

At the end of this blog, I’ll describe an Occam’s Razor, but for now let’s look at things a bit more management school style.

A             “Data Scientist”'s Point of view


The top right is a sweet spot for the “Data Scientist”; he/she gets to continually learn things and use the skills in a variety of projects. This is a happily tired “Data Scientist” where it’s not work, just play.


The bottom right is where the “Data Scientist” is kept busy on projects that require his/her skills and knowledge, but these projects are repetitive. This can lead to boredom, and the palliative situation is to build strong feedback loops to keep improving and challenging the “Data Scientist”. Else it might make sense to rotate “data scientists’ and bring in new pairs of eyes to try and bring quantum improvements rather than continual polishing.

The top left is where the “Data scientist” gets a variety of projects but is under-utilised, a stop-start adventure. This is likely the case of an organisation who is starting with data science operationalisation, or is not mature enough to exploit “data science” fully. In this case, it might be better to hire specialist “data scientist” on a as-needed basis. This would be much better use of resources and make better use of “data scientists’” skills.

The bottom left is where no “data scientist” would want to tread; apparently he/she did not ask the right questions in the interview.

B             Hiring Manager's Point of View:

I was going to write an analysis of a hiring manager’s point of view, but I think it is enough to say that if they are not aligned to that of the organisation (principal agent problem) then it doesn’t matter what is really going on, all that matters is the impression you can give; it is an ego or resume padding trip and no reasoning can be applied.

C             Organisation’s point of view
The top-right is the sweet spot for the organisation; the “data scientist” is involved in many projects (well utilised) and a variety of projects (organisational analytical maturity), chances are organisations in this quadrant are able to generate sufficient RoI (Returns on Investment) from their “data scientist”. But does that mean that organisations in this quadrant should not hire data scientists on gigs? The answer depends what the “data scientist” can bring.

The bottom right is where the “data scientist” keeps doing the same things. Basically, at every iteration of a model, unless there have been structural changes, there will be incremental improvements. The question is whether this generates sufficient RoI for an FTE.

One of the questions clients (both when I was FTE and on gigs) often ask me is when they know a model needs to be relooked at. My usual answer is all to do with metrics. In the same way as I insist of proper performance metrics of models to be agreed at the start of a project (to ensure RoI), I also encourage clients to have acceptable variability in results. It’s a bit like having a point estimate and a band of acceptable intervals (CI). So only if the results degrade below an acceptable level would it be worth looking at (you cannot assume that the degradation is due to random fluctuations and the impact on returns is too negative).

So chances are, unless each application of the model generates huge returns, it might make more sense not to have a “data scientist” on the payroll but to use a “data scientist” on gigs (1).

The top left is where the “data scientist” is engaged on a variety of projects but not many projects. This is a clear case that a “data scientist” by gig is better. Not only do you utilise resources only when you need them, but you can ensure you get the best resource for the project.

The bottom left shows and organisation that has no use for a full time data scientist, but should explore “data science” via gigs.

Conclusion:

It is quite clear that the best case scenario for both the “data scientist” and the organisation  is when the “Data scientist” is fully engaged, building a variety of models/algorithms to solve different business use cases and generating RoI.

In the rest of the cases, the situation is unlikely to last. In cases where there is either variety or quantity but not both, the “data scientist” is likely to feel lack of growth and leave, causing the organisation to go through the expensive hiring process again and again, further impacting RoI.

In the case when there is neither quantity nor variety, there is no point engaging a full time “data scientist”. This is a very clear-cut case for having “data scientists” on a gig basis. The approach should be one of proof of concept, prove the RoI that can be obtained, using “data scientists” on gigs; not only is the cost overall lower, but the organisation can get specialists.

Simlarly, when there is variety but lack of quantity, “data scientists” on gigs offer the possibility of specialist help, and help only when needed. When you do not have projects requiring skills of a “data scientist” then why pay someone?

When there is no variety, the challenge is different. It is likely that this is the case when the analytical development of the organisation has stalled; for example, “data science” is used only in one area of the organisation, hence a lack of variety. Here again it might make more sense to only get “data scientists” on gig basis, to try and expand the variety and increase RoI by opening new avenues for returns. Also, another way of increasing RoI in this case if to hire “data scientists” for prototyping, but leaving maintenance to less expensive resources (2).

Finally, does that mean that if you have both variety and quantity you do not need “data scientists” on gigs? Well, it depends how well your RoI is doing. In the corporate world, where performance has to increase over time, using specialist help for prototyping, having a “different set of eyes” looking at business issues can be a solution. Of course you may choose to rotate your “data scientists” to ensure that freshness, but if could come at the cost of returns.

In a nutshell:

In a nutshell it all depends on the RoI. When I first took up a contract, it was very exciting as someone in the analytics field. My headcount was directly funded by the business and I had to justify my existence every year, I had specific RoI targets. That made me so alive. I guess that’s why I am believe strongly in RoI.

Apart of vanity and bragging rights on the part of management, why would an organisation spend on a resource that is generating low or no RoI?

Organisations have to measure the returns on their investments; that includes human capital. As long as the RoI is met, then hiring a FTE “data scientist” makes sense. Else, at the least, until the organisation matures enough to be able to allow the “data scientist” to generate that RoI, it should stick to “data scientists” on gigs.


1 There have been cases when clients want to put data scientists on a gig but also pay a retainer. This is sort of a compromise between the 2 models. However, retainers may not work. From the organisations perspective there will be an incentive to use the “data scientist” for non-“data science” work.
2 Another question I often get is “when do I need to review my models?”. I believe that the model metrics should not be decided just for a one time use, but also for on-going performance. For example, a targeted response rate of 15%, where review will take place if the response rate dips below say 12% for 2 consecutive runs.

Wednesday, 25 April 2018

Yes, facebook has taken liberties with the data they collect about you, but how safe is your DNA?




A while ago, I wrote about a new insurance product launched in Singapore that required you to submit your DNA as part of the deal – you got ‘personalised’ advice in exchange. The ad ridiculously showed two identical-looking twins receiving different advice (since identical twins share the same DNA...). (1) In that blog-post, I mentioned that the insurance company was at pains to stress that they had no access to the DNA, but I raised the prospect of someone buying that company collecting the DNA and not being bound by the same rules. And unfortunately this prospect is very real.

Let’s take a step back, am I talking about DNA or facebook?

These few weeks have been exciting for people interested in data and “Big Data”, since the extent of the data collected by Cambridge Analytica via facebook, very often without the subjects being aware (2). I have been going on about the need for us to own our data but this really takes the cake; you were not only giving away your data but that of your connections too (53 Australians took the test – and possibly gained something – but the data of 311,127 was harvested. Similarly 10 New Zealanders did so, and data from 63,724 as harvested. I am not saying there were national boundaries, but these numbers give an idea of the pandemic).

Ok, so people’s surfing habits, likes comments, photos they posted in public were accessed and used, but what use can be made of this data? As the time magazine article (1) mentioned, one use was for Mr Trump’s presidential campaign. And as this article shows, the efforts started in 2014 (4), and were very effective as confirmed by Mr Trump himself (5):
But they had this expression ‘drain the swamp.’ And I hated it, I thought it was so hokey. I said, ‘that is the hokiest, give me a break, I am embarrassed to say it.’ And I was in Florida where 25,000 people were going wild, and I said, ‘and we will drain the swamp’ — the place went crazy. I couldn’t believe it. And then the next speech I said it again and they went even crazier. ‘We will drain the swamp… we will drain the swamp,’ and every time I said it I got the biggest applause

So we can at least say that the data facebook ‘allowed’ Cambridge Analytica to harvest from the subjects was, at least, ‘useful’.

So what does that have to do with DNA?

Basically if you think that someone getting their hands on your surfing history and using it for their own purposes without your consent is bad, what if they get their hands on your DNA?

The organisation that holds the DNA for myDNA from Prudential is Prenetics Limited (7). Recently I read that Alibaba and Ping An insurance are the major investors in Prenetics (8). On one hand, I find it amusing that Ping An possibly have access to data that Prudential help collect. On the other I find it scary that the data of these people (of course I did not purchase myDNA) is now in the hands of another insurer.

Anyway, Prenetics claims that the DNA of over 200,000 people across South East Asia, China and Hong Kong were in their hands as early as October 2017 (9).

But, I am sure some nice people will say, there is a legitimate reason to do research into DNA; hospitals and universities have been doing so to the benefit of mankind for years. Yes, but I would argue that the CEOs of hospitals and universities have different experiences as compared to the CEO of Prenetics (Mr Danny Yeung) and that may affect how the data is being used:
Prenetics started out as ‘Multigene’ in 2009 when it span out from Hong Kong’s City University. Yeung joined the firm as CEO in 2014, after leaving Groupon following its acquisition of his Hong Kong startup uBuyiBuy, and it has been in startup mode since then. Prenetics has raised over $52 million from investors which, aside from Alibaba, include 500 Startups, Venturra Capital and Chinese insurance giant Ping An.”

This, I will admit, is pure speculation on my part. For all I know, Prenetics really wants to help mankind and bless everyone whose DNA they hold with better health and lower health care costs (prevention rather than cure). But I have other reasons to be sceptical.

Basically, even if humans ‘decoded’ the whole DNA sequence (which hasn’t been achieved yet (10)), even if you have inherited a predisposition to a condition, nobody can tell where you will actually get affected by it:
Genetic testing can provide only limited information about an inherited condition. The test often can't determine if a person will show symptoms of a disorder, how severe the symptoms will be, or whether the disorder will progress over time.” (11)

And to make things more interesting, the pieces of the genetic code that have not been sequences were considered useless or too hard to analyse given technological limitations, but are now being re-evaluated. Does that sound familiar? For people in the “Big Data” space (especially proponents of the “Data Lake”), it should.

One of the arguments of the “Data Lake” is that we do not know what data can be useful; even if we cannot extract is and use it now, we might as well keep it since it might be useful.
When I first started in this line of work, the kind of conversations I would have would be along these lines:
Q: “What data do you need?”
A: “Just give me what you have and I’ll analyse”
Q: “That is impossible, tell me what data do you need?”
A: “Ok, can I have the list of pieces of data that you have?”
Q: “That is impossible, tell me what you want and I will see if I have it...” ad nauseam

Now technology and acceptance of the usefulness of data have advanced and it is possible to “keep all the data” in a “Data Lake” or “Data Swamp” as some friends call it (Drain it! Drain it! Sorry I got caught for a moment).

Pieces of data that we would have had trouble analysing a few years ago such as weblogs, or pictures, or voice recordings can now be analysed relatively easily. But these pieces of data were routinely considered to be useless.

It is the same thing with DNA data. And to make it worse, there is the link between being at risk of some condition as per your DNA profile and actually getting that condition.

Basically, there is way too much data that would be needed to transform this ‘risk’ into something that can be measured with ‘enough accuracy’. That is what insurance companies try to do when they ask questions about your lifestyle, smoking, drinking... but these are very crude.

So is it fair that you could be penalised because of a feature of your DNA make-up? Are we slaves of our DNA?

What I am getting at is not the importance of DNA data, but rather at the care that must be taken when conclusions are made, and people penalised for things they may not be aware of.

To make things more fun, not only is Prenetics in China, Hong Kong and South East Asia, but it has recently acquired DNAFit (12). This impacts Prenetics in 2 ways. Firstly geographically, DNAFit’s market presence is mainly in Europe and is expanding to the USA. Secondly DNAFit goes direct to the consumer whereas Prenetics tended to reach the consumer via Insurance or Medical companies. (In fact even Linkedin is one of DNAFit’s customers).

The impact of direct-to-consumer DNAkits is debatable (13), but “a little learning is a dangerous thing” (14), add to this the emotional weight of ‘learning’ not necessary pleasant things about your own self...

So what I am saying is:
  1. As individuals we should have control over the data we produce by living (web/call/messaging behaviour, surveillance footage...
  2. But we should also have control over data we produce by existing (DNA).

I think there are many gaps between the general public (who have no issues with being facebook’s product in exchange for a quiz (15)) and those who have some idea of what can be done with such data; the same for DNA. And it is critical for people to be educated or educate themselves on this. As long as there is such an asymmetry of information, together with major issues with how people/machines use the data (people/machines, not technology or data itself), the cost of exploitation can be very high.

I would like to end this post with the poem by Alexander Pope (14):

A little learning is a dangerous thing ;
Drink deep, or taste not the Pierian spring :
There shallow draughts intoxicate the brain,
And drinking largely sobers us again.
Fired at first sight with what the Muse imparts,
In fearless youth we tempt the heights of Arts ;
While from the bounded level of our mind
Short views we take, nor see the lengths behind,
But, more advanced, behold with strange surprise
New distant scenes of endless science rise !
So pleased at first the towering Alps we try,
Mount o’er the vales, and seem to tread the sky ;
The eternal snows appear already past,
And the first clouds and mountains seem the last ;
But those attained, we tremble to survey
The growing labours of the lengthened way ;
The increasing prospect tires our wandering eyes,
Hills peep o’er hills, and Alps on Alps arise !


7 https://www.prudential.com.sg/en/prumydna/mydnapromotnc/ see point g: ““myDNA report” means the personalised report that Eligible Customers receive from Prenetics Limited”

Tuesday, 17 April 2018

#metoo, the question of identification, understanding and biases we and AI may not be aware of


I am more than halfway through a “data science” blog when I read a piece of news where a reporter was kissed on both cheeks during a report on the Hong Kong rugby sevens (1). Another thought piece mentioned that the reporter looked humiliated afterwards. (2)

What I found most interesting is that the headline called the men “rude”. Rude? Rude is not saying hello to people, not saying “thank you” or “please”. Is kissing someone, in an obviously pre-planned way, just rude? Especially given that the person was apparently humiliated, then it ought to be more than that, no? Where is the #metoo movement?

Then I recalled another recent incident on American Idol where a judge, hearing a contestant has never kissed anyone before because the contestant believed that this would require being in a relationship, tricked the contestant into a kiss on the lips. (3). Again, there was minor backlash, but no #metoo against the judge. (4)

When you compare this to the groundswell of the #metoo movement, you cannot help but wonder... The crux of #metoo is identification, you identify with something, someone... May be it’s not just sexual harassment, or even sexual harassment of a woman.

Earlier this year I read this open letter that I highly recommend to everyone to read. (5). Emma Watson talks about feminism and what I like most is the idea that introspection is needed to understand your own position, especially things you may not be aware of.

When I gave my UN speech in 2015, so much of what I said was about the idea that “being a feminist is simple!” Easy! No problem! I have since learned that being a feminist is more than a single choice or decision. It’s an interrogation of self. Every time I think I’ve peeled all the layers, there’s another layer to peel. But, I also understand that the most difficult journeys are often the most worthwhile. And that this process cannot be done at anyone else’s pace or speed.

When I heard myself being called a “white feminist” I didn’t understand (I suppose I proved their case in point). What was the need to define me — or anyone else for that matter — as a feminist by race? What did this mean? Was I being called racist? Was the feminist movement more fractured than I had understood? I began...panicking.

It would have been more useful to spend the time asking myself questions like: What are the ways I have benefited from being white? In what ways do I support and uphold a system that is structurally racist? How do my race, class and gender affect my perspective? There seemed to be many types of feminists and feminism. But instead of seeing these differences as divisive, I could have asked whether defining them was actually empowering and bringing about better understanding. But I didn’t know to ask these questions.

Basically, what I am trying to say is that the stories above (the reporter at the rugby sevens and the American idol story) didn’t create that huge an outcry possible because of some bias which people may or may not be aware of.

So how does that relate to “data science”?

Well, it seems blogs get tagged by keywords, and I would rather this doesn’t get tagged as a political post, I have to add a “data science” bit; and for this purpose I will use the words data science without quotation marks.. So here comes the data science bit... 

As data science gets more and more automated, as ML and AI become more popular (especially to non data scientists), one of the hidden dangers is bias.

You are what you eat, even if you are a machine or are artificial.

It takes a lot of effort to even identify biases from a bunch of data. That’s something I mentioned before in the context of Human Resource Analytics where the impact could arguably be the worst (6).

There have been many articles such as on whether computers can be racist (7); the answer is yes if the data set which was used to train is happened to have a bias towards or against  a certain race even if it was purely unintentional. For example if your area of the world has virtually no orange people and one happens to apply for a role (say president) and gets accepted, a machine could pick up that orangeness makes one suited for presidency (The probability of becoming president given the race is orange is 1 haha). 

Serious thought is being given to the topic, Barocas and Selbst (8)  argue that “Addressing the sources of this unintentional discrimination and remedying the corresponding deficiencies in the law will be difficult technically, difficult legally, and difficult politically. There are a number of practical limits to what can be accomplished computationally.” Ransbotham (9) from Boston College also argued that having more data doesn’t necessarily remove sampling bias. 

In sum, when we, whether as data scientists or normal human beings want to analyse and issue, it is a good idea to understand our own biases and the biases in the data we have, so that we can do justice to interpreting the information we have and generating results.

(3) https://www.youtube.com/watch?v=ce3_D3IG96w I guess people who didn't know of this case would have assumed the genders were reversed. While some men would have loved to be kissed by Katy Perry, not everyone would, especially people who believe you need to be in a relationship before you kiss someone.
(4) Just FYI the contestant did not get past the round.

Thursday, 12 April 2018

Taxing teachers and national defence personnel but subsidising private companies


Recently, it has been announced that teachers will be charged for parking at their place of work (schools) (1) and so will military personnel (2) (3). On the other hand, as I highlighted in a previous post (4), public space has been reserved for parking of privately owned and for-profit operated bike leasing companies.

I found it a bit strange.

Yes there is a move to decrease the number of cars and increasing the use of ‘greener’ transportation methods (electric cars, car-pooling but also bicycles)(5) but is that a reason? Do we want our teachers to cycle to work so they are fitter and set an example for their students and fight obesity in schools? 

Do we think that portly people in military attire are unsightly (6) and aren’t the regular incentivised programs (7) sufficient to tackle the problem?

Even if that was the case, then why, on the other hand, subsidise the commercial profit making businesses that own the bikes and play in the ‘bike-sharing’ space?

I know that for many people the word ‘subsidy’ is quasi-taboo, but I am not using it lightly. In fact, one of the reasons why teachers have to pay for parking at their places of work is precisely to ‘remove the subsidy’ that they had been enjoying:
"Such practices are tantamount to providing hidden subsidies for vehicle parking and are not in line with the requirements laid down in the Government Instruction Manuals," the AGO had said.(8). 

Saying teachers are being taxed therefore is an exaggeration, but to the person having to pay for something they did not have to pay is equivalent to a tax; bottom-line you have less to spend.

As for the military bases, the rule doesn’t apply across the board, but only to specially chosen bases where ““Due to their proximity to public amenities, the car parks in these camps are deemed to have market value,” the ministry said” (9). For teachers it is across the board.

Do I really mean that the money taken from the pockets of teachers and military personnel gets transferred to the pockets of the owners of the for-profit bike-sharing companies? Of course not.

But do I find that the policies, while each standing on their own arguments contradict each other? Yes. And that is my point.

Remember the fish-ball-stick incident? (10) It illustrated the fact that different government agencies were not coordinated and that coordination is now the job of the Municipal Services Office (MSO). 

Actually, interestingly, in one of my previous roles, we had proposed to the government agencies a system that would take the feedback received from the public and automatically distribute it to the right authority or combination of authorities to deal with. But then we were told the MSO was already on the way. (You see, analytics can even help clear fish-ball sticks, just attach a drone haha).

Ahum, so am I saying that this tax and subsidise issue could have been prevented, or at least highlighted to the relevant authorities?

To put it simply, yes. What I did was simple: I realised what the implications of different policies were, what space in economics they occupied, then I saw that their positions were in opposition not to say contradictory. Is it difficult to build a simple analytics based system to do that? No, it is not that complicated and there are quite a few algorithms that can help get the topics, stuff like LDA (11) or LSA (12), or a simple Bayesian classifier (13). Then it’s a question of comparing documents on the same topic.

“Data Science” to the rescue? Anyone in the government would like to know more? :D


3 Actually in the latter case the work places where the payment is being implemented has increased, not a totally new policy
5 Note that companies like grab (and uber) do not make the world greener, on the contrary. The price point of such companies is lower than regular taxis but higher than busses/trains which are more efficient means of transport. SO moving people away from taxis to grab does not make a huge green dent, but moving people from busses/trains to moves people to less green means of transport. That could be the topic of another blogpost... J
8 as (1), AGO means Audit-General’s Office
9 as (2) above, the ministry being the ministry of defence.