Thursday, 28 January 2016

Some thoughts on what Uber/GrabTaxi/lyft can easily do with your data




“Exclusive: The ride-sharing company is conducting a trial in Texas using movement sensors in smartphones to track signs of erratic driving” http://www.theguardian.com/technology/2016/jan/26/Uber-monitoring-drivers-us-passenger-safety-houston
When I read the title of the article I wanted to laugh… It would be really dumb of Uber if this hasn’t been already happening for a while.

Imagine, as an Uber driver, you carry the Uber app in your phone; do you control what the app accesses? Your smart phone has in-built GPS (that’s what allows you to use the maps and other applications that give you directions, recommend stuff ‘near you’), gyroscope (that’s how you can play games that require you to move the phone itself rather than just use controls). 

Using these 2 pieces of equipment that most smart phones have in-built, it is very easy to know where you are, the speed and direction you are moving, and any sudden changes in direction. All you have to do is overlay a map that has data on the speed limits, and you will easily be able to tell whether speed limits are being exceeded, whether sudden lanes are changing are taking place…

So if this is an exclusive story, then I am really surprised. Uber can’t be that backward, afterall, remember the “rides of glory” where Uber classified the rides their customers took? (For example if you are picked up in an area with healthy nightlife in the wee hours of the morning and get dropped off at a new location, then depart from that location after a few hours later, it could be a one night stand…)

You cannot be doing this and not analysing what your drivers are up to, especially after the horror stories such as http://indianexpress.com/topic/delhi-Uber-rape-case/ . In fact, wouldn’t it be easy to simply track every Uber car, since you know the origin and destination of every trip a passenger takes, decide the most likely routes, and issue alerts when driver deviate from these routes, may be require driver and passenger to respond?

What is more even interesting is what is going to happen/happen to this data that Uber collects. 

Data is a resource, and a very useful one too, in the right hands.

MSIG recently announced it would be the first insurance company in Singapore to use telematics (http://www.asiainsurancereview.com/News/View-NewsLetter-Article?id=34686&Type=eDaily) . To be able to decide whether a driver is low or high risk, the insurer will have to collect the data on how that person drives. But for Uber drivers, Uber already has all this data.

Take a step further, think of passengers. 

When you fill up your application for insurance, you have to declare your habits, including the amount of alcohol you consume, and whether you enjoy sky-diving for example. It is up to you to inform the insurer if the response to any of the questions you answered in the past has changed. Else you could lose coverage at the most inopportune times (http://www.mcmha.org/never-lie-life-insurance-application/). What if your insurer knows you are often at a well-known nightspot until the wee hours? Did you declare alcohol consumption? This can impact your premium or seriously cut any pay-out from the insurance company. The insurance company would mostly likely gladly pay for information that would save them from large pay-outs.

What if your bank knew you often went to the vicinity of the casinos and add it to the fact that the bank knows that you do not spend on shopping there? How would that affect your credit rating and your ability to get loans at decent rates? The bank would also be most likely to be happy to buy data that ensures their risk is well covered (high-risk customers are charged ‘appropriately’ higher premiums).

What if the data on trips taken by Uber is sold in bulk? Nowadays with the proliferation of data, it takes less than 5 matches to be at least 90% sure of one’s identity and thus know who you are, where you were, and when. http://www.nature.com/articles/srep01376. Basically, all I’d need is 4 social media posts where you confirm your presence at certain locations at specific times that I can match with the Uber data (starting point or end point and time), and I can identify that all the Uber trips are made by you with 90% accuracy at least. 

Basically what I am trying to say is that there is a lot that can be done with data that Uber (or any other car ride company – I am just talking about Uber because the article that caused me to write this was talking about an ‘innovation’ by Uber), whether on its own to enhance Uber services (is the driver a safe driver? Is there a risk that the driver is up to no good?), or be used by third parties (is this driver safe and deserve a low premium or unsafe and merit a higher one to cover my insurance risks?) or merged with other data to reveal even more (is this customer a low risk customer or does he/she have less than prudent financial habits? Where does Mr XXX or Ms YYY go, where was he/she last Friday night, where does he/she visit often?)…

And this brings me back to my main concern; do we have any control over what happens to data we generate? Laws are required to force organisations that collect data that we generate to inform us of and get our permission for data collection and usage, and well as retention period and delete data upon request.

Sunday, 13 December 2015

Think!




This sign on the left is found at a coffee shop in Singapore. A customer has a choice –after queuing for food – you queue again for drinks, or wait for an uncle/auntie who takes drinks orders and you pay upon delivery. The sign shows a productivity drive by the coffee shop owner.

Productivity can be defined as output per worker; so by encouraging customers to by-pass the uncles and aunties and going straight to queue themselves, there are less workers for roughly the same volume of drinks sold; abracadabra! Productivity goes up. The coffee shop owners might even gain even more: http://www.mom.gov.sg/newsroom/press-releases/2015/0819-leds

People who have lived in Singapore for a while will certainly understand that using the terms ‘productivity’ and ‘efficiency’ are words that push people to action. And there now is a long queue of customers at the drinks stall, and fewer uncles/aunties employed to collect orders.

As customers, what have we gained?

1.      Are the drinks cheaper? No.
2.       Is the waiting time for drinks shorter? No, on the contrary, instead of preparing 5 hot coffees,5 teas, and 3 milos the people manning the counter have to prepare them individually, increasing the time taken for each drink.


So we are paying the same amount for the drinks, only we are spending time queuing rather than spending the time with our friends and families and sharing a full meal together.

This is a simple example of an organization shifting costs to the end customer. The customer ‘pays’ more and the organization reaps savings. What savings? Some of the uncles/aunties are no longer seen at the coffee shop. I found one uncle, at a coffee shop in the next town, presumably further from his home.

So my question is, why do we as customers help these organisations fire people who obviously need the jobs?

I was having dinner with this friend last week and she was arguing that no one can stop progress automation…. Please, think!

This is not automation, no some soulless blind machine taking over the jobs of these uncles and aunties. We, as customers are actively taking hammers to the rice bowls of the uncles and aunties by reacting like Pavlovian dogs to the words ‘productivity’ and ‘efficiency’. We are choosing to take the extra effort upon ourselves to drive these people out of their livelihoods. And they might end up on the right side of the picture above, a cardboard auntie’s life is much harder than a drinks auntie’s life.

Think people, think!

I would have no problem paying an extra 5 or 10 cents per cup of coffee that would have meant the uncle/aunties could retain their jobs; I can’t be the only one right?

Thursday, 26 November 2015



Italian Wines and k-means

I had a couple of hours to kill, and given that I have some high class friends who requested a piece on wine rather than single malts (ok, some people will complain it’s Italian, but still…). I am trying to show how easily a simple old fashioned algorithm like k-means (it is about 50 years old) can do a decent job of classifying data. I am also trying to show that it is important to know what you are trying to do and choose the right algorithm and the right way to apply it.



The starting point is the wine dataset, a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars (you’ll have to taste and decide which is which, if you do please let me know), and there are 13 different dimensions: alcohol, malic acid, ash, alkalinity of ash, magnesium, total phenols, flavonoids, non-flavonoid phenols, proanthocyanins, colour intensity, hue, od280/od315 of diluted wines and proline. The dataset contains 178 data points classed into the 3 cultivars which I split into a training and validation dataset.

Basically I am trying to get k-means to learn what makes each cultivar (k-means works on centroids, hence the centroid can be thought as the ‘typical’ chemical composition of that cultivar) from the training dataset, then apply it to the validation dataset. The metric I will use for accuracy is simply the percentage of correctly classified cultivars.

First, I blindly applied k-means. We already know that the actual number of clusters in the distribution is 3, so I set K to 3 and use all the variables. 

This is what a simple visualisation of actual distribution of the validation dataset looks like; each node represents a wine type; they are colour coded as per the cultivar they belong to: red for cultivar 1, green for cultivar 2 and blue for cultivar 3. 



Keeping the original colour scheme, we can see that k-means doesn’t do a very good job at predicting the right cluster for each cultivar.




Predicted
A
B
C
Actual
1
12
0
12
2
0
17
4
3
0
12
5

The accuracy of the kmeans is only 47%.

But if we decide to work off normalised values, the picture becomes much clearer:



Even visually this looks like a picture with clearer segments. Applying k-means on this normalised data:





K-means does a much better job at distinguishing the cultivars based on the normalised data:


Predicted
A
B
C
Actual
1
24


2
4
1
16
3

17


The accuracy of k-means jumps to 92%.

In conclusion, it pays to have a clear idea of what the objective is and use the right approach. In this example, it should have been obvious right from the outset that standardisation is necessary; without it the results would be skewed by the fact that the scales for the various measures are very different and this difference distracts from the aim of the analysis. Also, we can see that there sometimes is no need for very sophisticated techniques, and a quick piece of analysis can deliver reasonable results with some fore-thought.
 
The original source of the data is:  https://archive.ics.uci.edu/ml/datasets/Wine