Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Friday, 13 September 2019

A rubbish clean-up

That rubbish set of stats over at the Keep New Zealand Beautiful website, noted earlier this week, is now corrected.

The National Litter Audit website has been purged of the bogus numbers, and the report updated.

This is good.

Unfortunately, in the absence of any more formal retraction or notice from them to the journalists that reported so credulously on the figures, we're unlikely to see either any correction or any updated stories noting the figures are wrong.

Monday, 28 January 2019

Afternoon update

The worthies of the afternoon closing of the browser tabs:

Monday, 27 February 2017

A stupid Newshub beat-up [updated]

Newshub today helped make Kiwis just a little bit stupider. But Ministers not knowing the underlying stats didn't help. [See update below though!]

To recap. The government was put on the spot about whether they're rorting tourism numbers. MSD will sometimes put people in hotels or motels as temporary emergency accommodation. Whether that happens too often relative to an ideal is a different question we'll leave to the side for now. Question at hand is whether that's inflating the tourism numbers. 

Tourism Minister Paula Bennett was asked whether the tourism numbers were wrong because of this. 

The correct answer is "MSD clients are a tiny fraction of overall hotel nights, so it really cannot affect the figures either way." 

Minister Bennett clearly didn't know what's going on in the underlying stats because she said that they aren't included because they're not tourists. Hotels don't know why guests are spending the night. They just report up to Stats how many nights they've provided. [Update - see below] Other non-tourists included in the figures:
  • A couple getting a room for a discreet encounter, who aren't tourists;
  • Someone who realises he is in no shape to drive home and would rather spend the night in the hotel rather than go home drunk in a cab;
  • Someone taking a night at a hotel after a row at home;
  • Someone renting a room as a meeting space;
  • Someone staying in a hotel room during some renovations, or before taking possession of a place they've just bought.
None of it matters. Why? There are almost 22 million domestic guest-nights per year in New Zealand hotels, and over 15 million international guest-nights. How do we know this? The tourism satellite accounts. Here's Table 8.


Neither the international guest nights nor the growth in international guest nights is likely to have been affected at all by MSD clients; they wouldn't have been reported as domestic visitors. It is unlikely that MSD clients have any material effect on the overall domestic guest nights either - it would be like thinking the water volume of Lake Taupo is overstated because nobody netted out the mass of fish in the lake. Yeah, there's fish, but it won't make much difference to the overall figures.

How much effect could it have had? The Newshub story reports 8,860 emergency housing grants in the last quarter of last year at a cost of $7.7 million. Let's say that those are all hotel room nights. Since they're emergency nights, they're not going to be getting "book ahead and save" rates. And they're also potentially riskier for the motellier. Let's say that the room rate is $100 per night but I'd think I'm erring on the low side there. That's (top end) then about 77,000 nights in that quarter. If the room rate is $200/night, then it's 38,500 nights. 

If that had persisted for the whole year, the total number of guest nights would still have rounded to 22 million - but the measured growth rate would have been a bit lower. But, again, would it matter? The government crows about international tourist numbers and guest-nights. Domestic doesn't get noticed as much. 

Prime Minister English noted "if they're counting them as tourists, they shouldn't be." It maybe wouldn't be that hard for MSD to tell Stats how many nights they've purchased and then have those netted from the tourism satellite accounts, but it's stupid hassle for no particularly good reason. And unless they do it all the years back, they're going to break the continuous data definition. 

So, some bottom lines:
  • There is a housing crisis;
  • The government is not fudging the tourism stats by including MSD clients in the tourism satellite accounts, and neither is Stats NZ;
  • It is stupid, and damaging, and unethical, to undermine trust in official statistics in this kind of Gotcha! attack on Ministers who cannot reasonably be expected to know what's in the definition of particular stats - and especially where it is inconceivable that whether or not it is included it would make a whit of difference to the measured tourist night numbers. 
  • I hope that the Statistics Minister, on advice from Stats NZ, would also have told Newshub that the 22 million nights context means that this would just be rounding-error stuff anyway. If he did, and Newshub didn't report that part, that would be worse for them. 
  • If we ever get to the point where MSD emergency grants could materially affect the domestic accommodation guest-night figures, we're going to need a bigger word than crisis to describe what's going on in New Zealand's housing situation.
UPDATE: MBIE's tourism estimates, like the monthly regional tourism estimates, don't use the accommodation survey figures anyway. So if they did use them, it wouldn't matter because the numbers are tiny. But they don't. In this evening's reader mailbag (haven't seen a source link yet):

The Accommodation Survey is produced by Statistics New Zealand monthly to provide information on short-term commercial accommodation activity at a regional and national level. This includes all people staying in commercial accommodation, not only tourists.
 
Domestic tourism is currently measured by visitor spending, which is not informed by the Accommodation Survey.
 
MBIE does not use the Accommodation Survey to produce key tourism data products, such as:
  • Monthly Regional Tourism Estimates
  • International Visitor Survey
  • New Zealand Tourism Forecasts.
People with emergency housing grants are not included in these statistics.
 
However, while it does not inform these products and measures, the Accommodation Survey is part of a suite of statistics that we use to understand the tourism market, both domestic and international. People with emergency housing grants make up a tiny percentage of the approximately 38 million visitor nights recorded annually in the Accommodation Survey.

Tuesday, 12 January 2016

Data integrity and fraud detection

Lindsay Mitchell raises an interesting problem.

In 2011, the Department of Labour matched HLFS employment survey data with benefit data and found:
About 40% of people on work-tested benefits may not be meeting their labour market obligations, as they appear to be either working too much or searching too little.
The backstory Lindsay provides is excellent - it looks like the Ministry tried to bury the paper, and it came out later accidentally. Lindsay only got it with the Ombudsman's intervention. Go read her whole post. And she gives one plausible non-evil reason why the Ministry might have wished to bury it:
But imagine a beneficiary reads or hears about how a survey they are being forced to participate in is being checked against their Work and Income records. For the welfare abuser, that would merely tip them off to lie more consistently to government departments.

Data-matching is being used increasingly but its effectiveness lies in keeping the public in the dark. There's an irony at work. Non-transparency is required to improve integrity of systems.

So ultimately that's where I find the most convincing rationale. But that leaves me with a dilemma.

As a long-time critic of the welfare system, the findings vindicate or illustrate my concerns about the rampant misuse of the system (which hurts genuine beneficiaries and the taxpayers funding it). Do I want to make a song and dance about these findings though, if the information acts to assist those with the worst motivations?
So long as no beneficiary actually was punished for truthfully answering an HLFS or HES survey, the odds of contaminating future survey responses are lower.

The most paranoid end of the distribution would expect that the government has been doing this forever and so always would have lied; the least paranoid end would either expect that the government weren't competent to actually match up records, or that Stats NZ wouldn't be lying about the uses to which their data is put. Without actual cases of "I know a guy who told the truth on the HES survey and *bam* lost his benefit", I wouldn't expect huge effects - but I have low confidence in that expectation.

But I think there's a way around it.

First, link up IRD and HES/HLFS and MSD data from last year through the IDI, along with whatever other administrative data seems useful. Use the IRD and HES/HLFS data to establish true cases of fraud. Use the rest of the data to get the correlates of fraudulent receipt. If the data allows for a reasonable predictive model, great! Save the parameters for next year. If not, abandon.

Then, if the predictive model had been decent, use next year's administrative data to forecast which recipients are at higher risk of fraudulent receipt - and have MSD follow up the higher risk cases. Drop from the sample anybody who was an HES respondent - it'll be a pretty small number anyway. You'll then be pinging those recipients who are similar to last year's fraud cases, but you won't be hitting anybody who was one of the survey respondents. After enough of a lag, bring the prior year HES respondents back in - their back-end data should have changed sufficiently that they won't perfectly predict any more, so they are not being punished for having answered truthfully. They're being audited if their characteristics are still very similar to those of high risk cases.

It won't be perfect - there'll always be some who'll lie on the surveys, just in case. But would there really be many who'd start lying because of this procedure?

Wednesday, 25 November 2015

Police muzzles?

There's a strong academic freedom case against the kinds of restrictions that the New Zealand Police, and other government agencies, put on the use of data. But it's not quite as clear-cut as it might seem.

Jarrod Gilbert is New Zealand's leading expert in crime and gangs. Between him and  Greg Newbold, there's not much the Canterbury sociology team don't know about crime. Gilbert's having trouble getting access to police data on crime that he needs for his research, though, because his research has meant he's spent a lot of time with gangs.

This part seemed especially concerning:
The degree of control the police sought over research findings and publications was more than trifling. The research contracts demand that a draft report be provided to police. If the results are deemed to be "negative" then the police will seek to "improve its outcomes". Both the intent and the language would have impressed George Orwell.
Researchers unprepared to yield and make changes face a clause stating the police "retain the sole right to veto any findings from release". In other words, if an academic study said something the police didn't like - or heaven forbid was in any way critical of the police - then the police could stop it being published.
These demands were supported by threats. The contracts state that police will "blacklist" the researchers and "any organisations connected to the project ... from access to any further police resources" if they don't abide by police wishes.
I worried about this kind of thing in Ministry of Health RFPs a while back.

But we have to balance it too against the following kind of scenario. Suppose that you're a government agency who's contracted with some researchers to do some work for you. Somewhere along the way, things go off the rails. The researchers' drafts don't look good, the statistical analysis doesn't line up, and it's starting to look like they're trying to grind an axe rather than provide you with a down-the-line assessment of results.

Worse, they've started shopping around their working draft at conferences [legit!] and putting out press releases about their working draft [uh-oh], touting their initial results, using the Ministry's funding as imprimatur, and calling for policy change based on it. You know they've missed a pile of important stuff in their analysis and that the results really aren't sound. What do you do?

In an ideal world, academic freedom prevails. The researcher presents their results, but the Ministry puts out contextual information noting what the researchers have missed and why the results are only tentative and preliminary. But there's a lot of risk that's then come in. The opposition might have latched onto the preliminary work and called the government cowards for not having changed policy, or, worse, bought out by "Big Industry Interests". Or, even worse, the preliminary wrong results conform to the Minister's priors and the Minister won't even let you put out a contradictory note because they've already tasked you with formulating policy based on it.

I'm not saying muzzle clauses are justifiable, but rather that this can be where Ministries are coming from.

It is interesting, though, that the University of Canterbury in general is happy to sign contracts for government funding that put strong muzzles on its researchers. I suppose the presumption is that government contracts are always wonderful and that the restrictions are always for the best in this the best of all possible worlds. On the other side, it doesn't seem to matter how rock-solid the academic freedom provisions are in an external funding arrangement if industry's involved - somebody's going to object to it. Again, the one-sided scepticism problem.

Tuesday, 24 November 2015

Cinderella men

I think this is the first time I've seen evolutionary biology featured in a Press piece on crime. The piece notes the disproportionate number of children in New Zealand killed by step-fathers.
Why stepfathers kill their lovers' small children but spare their own has troubled Canadian evolutionary psychologist Martin Daly for decades.

He and his late wife Margo Wilson founded the Cinderella theory in the 1980s, researching the deaths of 700 Canadian children.

What they found suggested the unconditional love a parent feels for a screaming child who has soiled their nappy, is not innate for a stepparent - and makes them more likely to lash out.

... Building on Darwin's theory of evolution, the relationship between the new man on the scene and his lover's child is forged by biological altruism, Daly and Wilson found.

That means humans, like other animals, are programmed to investing their time into reproducing their own genes - not someone else's - and sometimes that resentment becomes deadly.

..."My argument, in psychology - and it's the same with those other animals who engage in step-parenting - is the step-parent is doing it as a courting step.

"People love their own children more than they love someone else's child. That's not to say they don't love them... [but] generally, they're not going to throw themselves in front of a truck for them."

In the Canadian research, birth parents overwhelmingly smothered or shot their children, and a third of fathers committed murder-suicide.

Stepfathers usually beat children to death and just 1 in 67 killed themselves too, Daly and Wilson found.
They note that the New Zealand data for a proper test would be hard to come by (as is all NZ data about everything because of because).

Satoshi Kanazawa explained it this way:
In retrospect, this makes perfect sense. Parental love for children is evolutionarily conditional on the children’s ability to increase the parents’ reproductive success. Stepchildren do not carry any of the genes of the stepparents, so there is absolutely no evolutionary reason for stepparents to love, care for and invest in their stepchildren. Worse yet, any resources invested in stepchildren take away from investment that the stepparents could make in their own genetic children. So, in the cold, heartless calculus of evolutionary logic, it makes perfect sense for the stepfather to kill his stepchildren, so that his mate (the mother of the stepchildren) will only invest in their joint children, children whom the stepfather has had with the mother and who carry his genes. Only they can increase the stepfather’s reproductive success.
But he also cites some contrary evidence from Sweden suggesting that the background characteristics of stepdads do a lot of the work.
In their paper, Temrin et al. do not question that stepchildren are more likely to be killed and maimed by their stepfathers; they only question discriminative parental solicitude as the explanation for it. They point out, and empirically demonstrate with a small Swedish sample, that men who become stepfathers, by marrying women who already have children from previous unions with other men, are more likely to be criminal and violent to begin with. And Temrin et al. argue that their greater tendency toward criminality and violence, not their genetic unrelatedness, is the reason they are more likely to kill and injure their stepchildren.

Once again, in retrospect, this makes perfect sense. Divorced women with children are on average older, so they have lower mate value than younger women without children. Given choice, and all else equal, all men would prefer to marry younger women without children rather than older women with children with other men. The logic of assortative mating would suggest that women with lower mate value are more likely to mate with men with lower mate value. And, as I explain in an earlier post, men with lower mate value are more likely to be criminal and violent.
So a proper New Zealand study would want to correct for the stepfathers' ex ante characteristics. It would also then partially answer Jan Pryor's question of the theory, raised in the original Press piece:
The theory also did not explain whether solo-mothers living financially strained lifestyles were targeted by men who preyed upon them and their children, Pryor said.
I'm curious how much of the effect here works through the biological Cinderella story and how much works through the assortative mating dynamics at the lower tail of the distribution.

Add to the list of "open questions that could be answered by somebody with time to muck around in IDI applications and the Stats Data Lab", or "Masters theses waiting to be written".

Tuesday, 17 February 2015

Moderate drinking is still good for you

Last week brought lots of headlines about a new study claiming moderate drinking doesn't really provide health benefits.

I'd previously reviewed the evidence here, here and here.

The new paper claims to better adjust for sick-quitter confounds by using lifetime never-drinkers as comparison group, but this is hardly a new technique: Rimm and Moats 2007 notably restricted their sample to healthy people who exercised and who had good diets - they found strong protective effects for moderate drinkers as compared to abstainers, even if the NZ MoH wants to pretend otherwise.

Statistician David Spiegelhalter walks through the latest evidence. Basically, there was no power to their test because they had too few never-drinkers in their sample, so nothing came up as statistically significant. He concludes:
So a more appropriate headline would have been "Study supports a moderate protective effect of alcohol".
In summary, the study is grossly underpowered to convincingly prove a plausible protection, and they have committed the cardinal sin of saying that non-significance is the same as 'no effect' in a study lacking sufficient events, in this case, deaths in non-drinkers. Maybe epidemiological studies should include power calculations, which make sure there is a reasonable chance of detecting a plausible effect, and which became standard in clinical trials after too-small studies were being used to claim that drugs did not work.
This is a poor use of statistics, and I am surprised it got past the referees and into the journal. A recent analysis showed that exaggerated health stories in the media were not generally the fault of the journalists, but the press releases they had been fed. Rather ironically, the analysis appeared in the British Medical Journal.
And see Snowdon, here.

How long until New Zealand's MoH starts citing this as further evidence against the health benefits of moderate drinking?

The NZ Cancer Society has been pushing a "no safe level" message around alcohol and cancer risk. As best I understand the literature, cancer risk is always increasing in alcohol consumption. But unless you have a strong family history of cancer as opposed to other things, it ought to be all-source mortality you look to for medical risks. The protective effects of moderate consumption are well established. But there does seem to be some determination to downplay those overall effects while emphasizing the disorders worsened by alcohol.

Wednesday, 21 January 2015

Needs a diff-in-diff

I'd love to see somebody else head into the Stats Datalab and do a bit more digging into recent declines in teenage fertility rates.

Teen fertility rates are down but we aren't entirely sure why. The new report commissioned by the Social Policy Evaluation and Research Unit notes increased use of contraception but doesn't have any clear reason why contraception use has increased. They note the link between deprivation and higher birth rates; it would be interesting if they presented results sorted by income cohort.

The time path has a sharp drop from the 1970s through the early 80s, then a slow decline, then a sharp rise from '05 to '08, then a reasonable decline since '08.

I wonder whether changes under National restricting the generosity of benefits paid to single mothers who have an additional child while on benefit have had an effect on teen birth rates.

Somebody with DataLab access could check:

  • differential effects on teenage fertility of the changes to benefits by comparing cohorts likely to access benefits conditional on childbirth with higher-income cohorts unlikely to do so;
  • differential effects of easier morning-after pill access in Auckland, later rolled out to other cities, as compared to regions where it's more difficult to access pharmacies;
Please go and make this your thesis and report back.

Tuesday, 15 July 2014

Stat Juking revisited

I'd reckoned you'd need a bit of stats-fu to find evidence of police juking of the crime statistics. Turns out there was an easier way. Bevan Hurley reports that the Herald on Sunday got a copy of a report showing that Counties Manukau police had been fiddling the burglary numbers by recoding burglaries as less serious offences. 
About 700 burglaries were “recoded” in the Counties Manukau south area over three years, an internal police investigation has found. It found that about 70 per cent of the time, the offences should have remained burglaries.
The revelations will be an embarrassment for Police Commissioner Mike Bush, who was district commander of the area at the time, although he was not responsible for overseeing the coding.
Police have not said why the statistics were altered, but say staff were not under instruction to do so. Tolley denied police were under political pressure to reduce burglary statistics.
You don't need overt political pressure to get this kind of outcome, just KPIs with strong enough incentives. On the plus side, they were caught. On the down side, I can't see how lower level staff doing the coding would have any incentive to muck the stats around unless they were getting pushed by those whose KPIs did provide such incentive. It would be really interesting to read the full report.
The review listed dozens of examples where break-ins and attempted burglaries were downgraded, including one case where police failed to follow up after a witness gave them a burglar’s registration number.
The review found the burglary recoding rates in Counties Manukau south at the time were 15 per cent to 30 per cent whereas other areas typically recoded about 5 per cent.
So where last week's rumours were about failing to pursue charges, which wouldn't have mattered for stats based on recorded complaints, downgrading the complaints to less serious offences would matter.

I'd be curious to know what kinds of lesser offences were artificially inflated to keep the burglary numbers down.

I hope that the Police stats units have informed any researchers who'd been using the incorrect figures of the updated and corrected series. Anything that relied too heavily on 2009-2012 Manukau data is now going to have to be re-done.

The Herald on Sunday broke the story on the 13th. Their version is gated. The Stuff version, which notes "It was reported" rather than crediting the Herald, is here.

Wednesday, 9 July 2014

Stat Juking?

Labour claims that National's instructed the police to charge fewer people to meet crime reduction targets.  HT: NoRightTurn
“Front line police and others in the criminal justice system are telling us police have had pressure put on by senior officers to reduce the number of charges they lay to meet the Government's targets,” Justice spokesperson Andrew Little says.
“Police are increasingly using pre-charge warnings as a device to not proceed with charges. At the same time I have heard of people being told to gather evidence themselves before police will consider bringing charges.
Labour’s Police spokesperson Jacinda Ardern said New Zealanders were owed an explanation.
“We won’t stop the cycle of repeat offending against women and children by lowering the threshold for prosecution.
“This directive coincides with a significant drop in the number of family violence prosecutions, while at the same time the number of family violence investigations has soared.
“It doesn’t help that police are still not recording domestic violence offences separately, or that access to many programmes aimed at stemming family violence are contingent on a prosecution.”
If it were true, how could we tell?

First off, the greater use of pre-charge warnings isn't a secret. It was something advertised as a deliberate move to free up police resources rather than bring charges for minor offences. Here's the Herald from 2012; here's 3 News from 2010.

The Police's fact sheet on Pre-Charge Warnings notes that they're only used for relatively minor offences:
Operationally, PCWs are more suited in urban rather than rural areas, and in particular the centres of larger localities with prevalent disorder or alcohol-related offending. The initiative targets those 17-30 years of age (highest rates of overall offending are in this age group, including for the key offences eligible for PCWs). Of the top five offences where PCWs are most commonly used, four have no victims (and are “Police initiated”, such as Possession of Cannabis or Disorderly Behaviour). Shoplifting under $500 is the only offence often resolved with a PCW with a recorded victim. 
That doesn't mean that there isn't juking, just that there mightn't be a presumption of juking. What would juking look like?

  • There has to be some optimal use of PCWs rather than taking offenders through the courts: low-level offences where a scare should be enough. If the rate of subsequent offending among PCW offenders were higher than the rates among those formally charged, this could be suggestive of too many offenders going through PCW, and especially if the re-offending rates were increasing with increased use of PCWs without subsequent ratcheting back of PCWs by police.
  • For non-PCW areas like family violence, we'd expect to see an increasing divergence between crime rates as measured by charged offences and crime rates as measured by survey responses to questions like "Have you been a victim of crime in the last six months". 
    • There are crime victimisation questions in the NZ GSS, but it's only updated every two years. 
    • The Ministry of Justice maintains the NZ Crime and Safety Survey, but the last iteration of it was 2009
    • It isn't survey data, but if the hospitals maintain data on the source of ED-presented injuries, you could look for a growing divergence in the number of assaults backed out of that kind of source and the police charge rate. 
  • In the absence of frequent survey data on crime victimisation rates, you might look to see whether policing districts with higher ex ante crime rates had increased conviction rates with declining charge rates. If the police were under pressure to reduce the number of charged offenders to keep the number of offences down, you'd hope they'd at least decide to avoid pursuing the cases that were least likely to yield convictions. If these pressures were then different across policing districts because of different crime rates, I'd expect that:
    • A greater proportion of charges in juked districts fall on repeat offenders rather than first-time offenders;
    • A greater proportion of charges in juked districts proceed to conviction as fewer of the less-certain cases get pursued.
    • I'd expect some action in the time path, like districts getting close to some target crime rate start slowing their charge rate more quickly than we'd expect from mean reversion.
I'd be pretty surprised if the police here were juking the stats: that Little's presenting the use of PCWs as evidence of juking, when it was rather well announced policy, doesn't give me confidence. I've not looked at all at the kinds of statistics suggested above. But it's what I'd expect somebody laying accusations of stat-juking to be presenting.

UPDATE: Farrar notes that it shouldn't even be possible to juke the stats by failing to charge as the main crime stat series is based on recorded complaints. I'd assumed that Labour was effectively alleging that the police were juking things by failing to record complaints which didn't proceed to charge, which was part of my "I'd be pretty surprised" prior.

Friday, 15 February 2013

Indicators

The Canterbury Earthquake Recovery Agency has been putting out a nice compilation of Canterbury-related stats for a few months.

In the October 2012 quarterly report, we find that growth in weekly rental prices in Christchurch has outpaced that in the rest of the country; rents here began from rather below the national average and are now above it. I wonder what would happen were some of those rental prices to be quality adjusted; I'd also love to see data on how easy it is for would-be renters to find accommodation.

We also see that there are far more skilled vacancies here than elsewhere; more recent figures I believe have Canterbury's unemployment rate below that in the rest of the country. I wonder to what extent lack of rental availability prevents inbound migration by those who could work on the rebuild.

I've been appointed to the external review panel for CERA's economic indicators. Academics wishing that other data were available, or presented differently, please drop me a note; I can compile useful suggestions for the next meeting.

Sunday, 27 May 2012

Say's Law of Humbug

Demand for that which cannot be done brings forth supply of charlatans. Baum knew it:
Oz, left to himself, smiled to think of his success in giving the Scarecrow and the Tin Woodman and the Lion exactly what they thought they wanted. "How can I help being a humbug," he said, "when all these people make me do things that everybody knows can't be done?
The Munchkins had a latent demand for humbug satisfied by the entrepreneurial Oz.

Chris Dillow points out a nice modern example: demand for expert forecasts. Subjects in Powdthavee and Riyanto's experiment were run through "The System" - a classic scam where you send a random set of stock market or horse betting predictions, toss from the set anyone to whom you sent the wrong prediction, do it again, then offer to continue sending predictions (for pay) to folks who received a few lucky hits in a row. What happened in the lab? Says Dillow:
And here's the thing. Subjects who saw just two correct predictions were 15 percentage points more likely to buy a prediction for the third toss than subjects who got a right and wrong prediction in the earlier rounds. Subjects who saw four successive correct tips were 28 percentage points more likely to buy the prediction for the fifth round.
This tells us that even intelligent and numerate people are quick to misperceive randomness and to pay for an expertise that doesn't exist; the subjects included students of sciences, engineering and accounting. The authors say:   
Observations of a short streak of successful predictions of a truly random event are sufficient to generate a significant belief in the hot hand.
It's easy to believe that this happens in real life. For example, the people who are thought to have predicted the financial crisis of 2008 are invested with an expertise which they might not really have.
The paper's excellent title? "Why do people pay for useless advice?"

I wonder whether basic training at high school in financial literacy and classic scams might do any good. But it's hard to overcome the demand for humbug. And the paper finds that student subjects with more correct answers in a statistical test didn't spend less on predictions. If these were the results for college students, how awful would a general sample look?

Monday, 5 March 2012

Promotional regressions

A mathematical equation for holiday happiness, or banal tautology?

Expedia hired a couple of researchers to analyze a survey about Kiwis' recent holiday experiences. It looks like they set up a regression with "how happy did your last holiday make you" on the left hand side and a bunch of stuff on the right hand side like "the weather was good", "there were plenty of activities", "it was value for money", "plenty of partying", and that "it made my neighbours envious"*. Then they put out a press release presenting the regression equation and coefficients as being the mathematical formula for holiday happiness.

And it's gotten press. Radio New Zealand called me for comment on it last night;** I told them that if I understood the method correctly, it doesn't exactly provide would-be travellers with useful advice for happier travels. Who doesn't look for good weather when on holiday? Worse, most of what's on the right hand side seems to be just alternative ways of expressing what's on the left hand side; it would be surprising to find somebody who reported having had a bad time on holiday but that it was good value for money. Value-for-money kinda has happiness as the numerator.

And there's reverse causality all over the place: are people flying more than seven hours away happier because of the more remote destination, because they're richer and can afford long-haul travel, or because people are more willing to spend more on stuff once they've travelled that distance than when they've driven out to the bach? Are couples travelling together happier because they're travelling in pairs, or because single people travelling alone tend to be unhappier people in general?

Here's the funniest part of the press release:

A group of experts including a psychologist, a mathematician and travel experts from
Expedia.co.nz set out to find just that, and have created the formula for holiday happiness.

With Kiwis taking over two million holidays a year, the formula proves there is a science that every holiday‐maker can apply to their next escape to guarantee themselves a top trip.

What are those sciency looking things? I'm pretty sure it's the following: HH is Holiday Happiness. It's a function of some constant {b}, "Good Weather" {GW}, "Plenty of Activities" {PA}, "returning home relaxed and de-stressed" {RD}, going to a Great Destination {GD}, finding Value for Money {VM}, engaging in Plenty of Partying {PP}, and inducing "Holiday Envy" {HE}. So update your vacation plans accordingly. Never mind that without units on any of the regressors, you've no clue whether the relative coefficient magnitudes signify anything, so you don't really know the tradeoffs on the happiness production frontier between "Plenty of Partying" and "Great Destination". Or that pervasive potential endogeneity problems mean you really can't use it as a guide for action.

At least Radio NZ looked for folks to comment on it - the reporter there noted having talked also with a statistician. Stuff.co.nz seems mostly to have copied the press release.
Psychologist Meredith Fuller, who worked on the equation, said people could use the factors when planning their next break.

"It's essentially taking the guesswork out of a holiday that could be an expensive disaster."

Choosing a top destination, getting value for money and lots of partying also feature highly.

The most satisfied people spent at least five nights away, took a long-haul flight to their destination, and travelled with between two and five friends. Fuller says five was the ideal number.

"There is less chance of individual conflict, more mental stimulation, and enough variety in interests to broaden our experiences."

A partner is fine, but including in-laws or family decreases the chance of holiday joy. "The worst people to go away with are your relatives," she said.

Lording it over your mates – more politely called "creating holiday envy" – also increases the prospect of a good result. About half of those surveyed used social media while away to post snaps to make friends jealous.

"It seems it's no longer a case of `wish you were here', but `I'm here and you're not'," Fuller said.
So, a lesson for the kiddies out there. Take your intro to econometrics course. You too can easily produce the mathematical formula for anything, at least as far as some of the media is concerned. All you need is r-e-g.***

And kudos to Expedia for getting press mileage on this one.

* Weird Al said it best, and without running a regression, in his ode to the world's greatest tourist destination:****
Then we went to the gift shop and stood in line
Bought a souvenir miniature ball of twine
Some window decals, and anything else they'd sell us.
And we bought a couple post cards, "Greetings from the twine ball, wish you were here!"
Won't the folks back home be jealous.
Note that I'm not at all disputing that people enjoy vacations. I just can't really see how the survey and regressions really provide any kind of useful guide to action.

** I have no clue whether anything aired. UPDATE: here. But they left out all the "these econometrics are crap" bits. Ah well.

*** Maybe they used ordered probit. I haven't the full paper.

**** I regret, but Susan does not, that we drove within an hour of the place but I didn't realize we were anywhere near it. Had I realized, we'd have gone. Missed opportunities.

Friday, 17 February 2012

Trusting econometrics

One of my profs at Mason told the story of how he'd been offered a new boat if he could get the coefficient in a regression to be below two - which would have allowed a merger to proceed. He turned it down, but not everybody does. Unfortunately, in a whole pile of empirical work, you either have to really trust the guy doing the study, or make sure that his data's available for anybody to run robustness checks, or check that a bunch of people have found kinda the same thing. Degrees of freedom available in setting the specifications can sometimes let you pick your conclusion, like getting a coefficient that hits the right parameter value or the right t-stat.

David Levy and Susan Feigenbaum worried a lot about this in "The technological obsolescence of Scientific Fraud". Where investigators have preferences over outcomes, it's possible to achieve those outcomes through appropriate use of identifying restrictions or method - especially since there are lots of line calls in which techniques to use in different cases. They note that outright fraud makes results non-replicable while biased research winds up instead being fragile - the relationships break down when people change the set of covariates, or the time period, or the technique.

Note that none of this has to come through financial corruption either: simple publish-or-perish incentives are enough where journals are more interested in findings of significant than of insignificant results; DeLong and Lang jumped up and down about this twenty years ago. Ed Leamer made similar points even earlier (recent podcast). And then there's all the work by McCloskey.

Thomas Lumley today points to a nice piece in Psychological Science demonstrating the point.
In this article, we accomplish two things. First, we show that despite empirical psychologists’ nominal endorsement of a low rate of false-positive findings ( ≤ .05), flexibility in data collection, analysis, and reporting dramatically increases actual false-positive rates. In many cases, a researcher is more likely to falsely find evidence that an effect exists than to correctly find evidence that it does not. We present computer simulations and a pair of actual experiments that demonstrate how unacceptably easy it is to accumulate (and report) statistically significant evidence for a false hypothesis. Second, we suggest a simple, low-cost, and straightforwardly effective disclosure-based solution to this problem. The solution involves six concrete requirements for authors and four guidelines for reviewers, all of which impose a minimal burden on the publication process.
Degrees of freedom available to the researcher make it "unacceptably easy to publish "statistically significant" evidence consistent with any hypothesis." They demonstrate it by proving statistically that hearing "When I'm Sixty-Four" rather than a control song made people a year-and-a-half younger.

The lesson isn't radical skepticism of all statistical results, but rather a caution against overreliance on any one finding and an argument for discounting findings coming from folks whose work proves to be fragile.

Wednesday, 1 February 2012

Bogus polls

Web polls are worse than useless, says Thomas Lumley. Why? Anchoring bias.
Seeing the results is likely to make your beliefs less accurate, even if you know the information content is effectively zero.
It might not be immediately obvious how bogus web polls cause harm. But if anchoring bias feeds into conformity or bandwagon effects, and that cycles into voter policy demands, we can move from bogus poll to "most people think X" to "Policy should be X" to "How can you oppose X, most people agree...". I think similar mechanisms work in bogus "cost of X" studies, eroding some voters' default liberalism by convincing them that they're bearing, through the tax system, costs actually borne by those engaging in the activity.

StatsChat continues in its Sisyphean quest to beat the stupid out of journalistic use of stats in New Zealand. Check the link above for fun and game in margins of error across three web polls. Forfty percent of Kiwis know these stats are bogus; shame it isn't eighnty.

Thursday, 12 January 2012

Memory in journalism

Keith Ng thoroughly documents one instance of a broader phenomenon in New Zealand journalism: very poor apparent institutional memory.

Keith notes that ACR, the Association of Community Retailers, which lobbies against regulations that impose costs on small tobacco retailers like dairies, gets PR support from Imperial Tobacco. Keith calls it astroturfing. That's possible, but I don't think it rules out that the group of represented small retailers genuinely supports the policies promoted by ACR and simply shares interests with Imperial on those issues. 

Either way, it didn't take long after Keith broke the story for the NZ media to forget that ACR enjoyed industry support. Writes Keith:
It seemed like quite a problem, the idea that so much of our news comes from groups which could be hiding all kinds of interests and agendas.
Turns out, the problem isn’t “what are they hiding?”, but “does anyone give a shit?”.
What this has shown is that even when the agenda is Big Tobacco’s, even when the connection is the second result on a Google search, even when their own organisation has reported on it, even when it’s stated plainly on their website, even then, the PR industry can get their stories printed with no scrutiny.
It’s a complete and utter rout.
You know, some of my best friends are journalists. And I like to think that they are the better ones. They always complain that they’re under pressure and under-resourced, and that this sort of shit slips through the cracks.
I’m sure it’s true, but here we are, at the point where our biggest news organisations run stories without spending 10 seconds on a Google search, or asking if something makes any goddamn sense. [Emphasis added]
If Richard Green says new laws will cost him $10k for shelves, they run it. If Richard Green says new laws will cost him $27k for shelves, they run it (RNZ newswire, 14 July 2011).
SHELVES.
How many of the facts reported in our media are this dodgy? And if there is so much that we can’t trust – and we can’t distinguish between what can and cannot be trusted – at what point should we simply give up?
You could chalk it all up to that it takes time to build up experience on a file and that erosion of profits in journalism have knocked out the senior folks who remembered what happened a year ago. But that can't be it when a 10-second Google search, or simply "asking if something makes any goddamn sense" would be enough to shoot something down. But when Radio New Zealand happily reports that each smoker costs the economy three times per capita GDP, folks aren't running the simple checks.

It has to come down to demand. If your audience take stats as infotainment, why worry too much if the stats are right? In fact, you can't afford to worry about it too much. And if the customers don't care whether the stats are right, then supply will arise to fill demand for the kinds of stats for which somebody's willing to pay (BERL on alcohol, PWC on Adult and Continuing Ed, many others in the back pages of the blog, the whole InfoGraphics problem cited by Megan McArdle...).

On bad stats, at least we have StatsChat. The University of Auckland's stats department now gives a prize for picking the week's (or month's) worst stat that's appeared in NZ press outlets. I recently nominated Radio NZ's exaggeration of the costs of smoking. Hopefully, shaming news outlets and the producers of bad stats will eventually have some effect. But it's harder to think of useful interventions that fix the kind of sloppiness Keith's citing.

Keith's interview on Radio NZ's interesting; it'll likely here be archived soon.

Tuesday, 13 September 2011

Lies and statistics

Here's the presentation I gave for reporters and others at the Christchurch Press a fortnight ago.


You'll probably need to put that on full screen for the small print to be legible.

Turnout wasn't huge, but did include reporters on the health and crime beats.

I recommended that the journalists start following StatChat. Here's hoping it all has done some good.

Friday, 9 September 2011

This week's nomination

I'm a fan of Auckland's Stats Department's "Bad Stat of the week" contest. I won their inaugural one by nominating John Pagani. I would have won another one but for lack of competition; alas.

Today's entry? Jennie Connor's trashing of a rather nicely conducted study on the health benefits of moderate drinking. Let's start with the study.

Qi Sun and a team including Eric Rimm found that moderate alcohol consumption correlates with good aging and health outcomes in middle-aged nurses in the US: the Nurses' Health Study. They follow a set of over a hundred thousand female registered nurses, starting in 1976 and surveyed every two years subsequently. It was set up to look at the long-term effects of oral contraceptive use but has been used in lots of other studies since.

They averaged reported alcohol intake in the 1980 and 1984 follow-up studies to get a measure of mid-life alcohol consumption, then checked to see whether alcohol consumption at mid-life increased the likelihood of successful aging, defined as being free of major chronic diseases and having no major cognitive or physical impairment and no mental illness.

To avoid the "sick quitter" confound, they dropped from the sample any participants who had:

  • Chronic diseases at baseline
  • Those diagnosed with alcohol dependence or chronic liver cirrhosis
  • Those reporting a significant reduction in alcohol use within 10 years before baseline
So, we shouldn't have former drinkers who quit because of poor health included in the reference group.

Next, they worried about confounding through other health behaviours. They controlled for a wide range of dietary, health, physical activity variables and life history of smoking. When you have observations every two years on a large group of women over a thirty year period, you can do a lot. But, you still can't completely correct for things you can't observe.

What can you do? You can check to see how much of the effect is reduced when you add in all the health behaviours for which you can control. If the effect drops substantially, it's likely there are other things for which you can't control that are correlated with the things for which you can control and might further reduce things. If the effect doesn't move much, it's not likely that other health behaviours are driving things.

Table 2, above, gives the relevant comparisons. Compared to non-drinkers, the odds ratio for those consuming 1.5-3 standard drinks per day was 1.26 in a model adjusting only for age and 1.28 in a model adjusting for all the health and activity correlates. It's exceedingly unlikely that unobserved health behaviours would attenuate the effect if correcting for observed health behaviours accentuates the effect.

They also run things within a cohort restricted to never-smokers. The pattern remains, though statistical significance is lost (you don't get a lot of never-smokers among nurses in a cohort that starts in 1976).

Ok. so what does Connor bang on about then?

The idea that moderate drinking may be good for health is well entrenched, but has been increasingly questioned, including by New Zealand epidemiologists Professors Jennie Connor and Rod Jackson, who challenged the theory in Britain's Lancet journal in 2005.
Professor Connor, of Otago University, told the Science Media Centre the Harvard study added nothing to the many similar studies that were "unreliable for answering questions about the health effects of drinking because of their design".
The supposed health benefit might be due to differences in lifestyle - other than drinking - that were associated with being a low-risk drinker.
"It may be true that women who drink one drink a day are healthier than others, but we do not know if it has anything to do with the alcohol, as these women are not the same as others in a variety of ways."
There was no scientific justification for the promotion of alcohol as health-enhancing for any sub-group of the population.
"The potential for harm is great, and the potential for good is unknown."

What had Connor and Jackson warned about in the Lancet? Two things that Sun, Rimm et al do a pretty good job in ruling out: confounding of never-drinkers with former drinkers, and unobserved health behaviours.

Jennie: if it's unobserved health correlates that are doing the job, why the heck does the relationship get stronger once they corrected for other health behaviours? If some underlying "healthy type" is driving things, that'll correlate with the other health behaviours that we can observe and will reduce the effect of alcohol on successful aging, not increase it. In other specifications, there are small reductions in the odds ratio in the multivariate model as compared to the bivariate one. But the reduction is small, like from 1.43 to 1.35. If throwing in a kitchen sink of health behaviours that are likely to correlate with the unobserved health behaviours reduces the effect by less than twenty percent, it's pretty unlikely that full controls would reduce the odds ratio to 1 or less.

And so I'm nominating Jennie Connor for the bad stat of the week contest for her trashing of what seems a pretty solid study.

Tuesday, 23 August 2011

Sports and science journalism

Thomas Lumley's fed up with lazy reporting in the sciences:
if a press release or a wire service story told you that the Wallabies had a new training regimen that would improve their game without making them fitter, faster, tougher in the scrum, more accurate with kicking, or better at putting in the elbow, you’d ask questions. We’d like to see science journalism eventually get up to the standard of sports journalism.
He points the finger at journos who gulped down the claim that an hour of TV watching costs 22 minutes of life expectancy [which I'd critiqued here]. And, because his stats-fu is better than mine (and probably because he read the studies a bit more closely than I did), he notes:
The journal, in its blog, says that the real story was about sedentary lifestyles vs exercise.  They are shocked — shocked — to find that there might be sensational news coverage of the article. However, they do note:“The blast of news coverage also suggests that creative research angles on behavioral health impacts are useful in grabbing the public imagination.” Indeed.
The really strange thing about the paper, though, is the 22 minutes of life lost per hour of TV.  Looking at the confidence intervals shows that there is huge uncertainty (the range is from 20 seconds to 45 minutes), but it still doesn’t really make sense.  Some of this is just correlation vs causation — in reality, not everyone who spends the evening in front of the TV would spend the time jogging, or even playing golf, if their Sky subscription were cut off.   The other component is the model used for years-of-life lost.  The researchers didn’t actually do anything with individual participant data from the AusDiab study. They took the results of a previous analysis (which didn’t get nearly as much coverage) and added in the assumption that the effect of TV weakened with age.  Under that assumption, the effect must be a lot larger for younger people than it appears, and since younger people have more minutes of life to lose, that increases the average cost of TV.  The assumption was described  by the authors as if it was a universal fact, and it’s true that several important risk factors follow this pattern — one reason is that there are more things to die of at older ages, which applies here, but another is that, eg, low blood pressure in older people happens for bad reasons as well as good, and that doesn’t apply here.  In any case the attenuation assumption is doing a lot of the work, and it is only weakly connected to any actual data.
I'm very pleased that there's somebody else out there who cares about this stuff. It's really a bit pointless though. Papers exist to sell eyeballs to advertisers. Most of those eyeballs are in front of squishy grey stuff that doesn't care a whole lot about scientific accuracy but rather about group affiliation and sensationalism. The real problem is on the demand side. We get the journalism we deserve. Just like we get the politics we deserve.

Saturday, 20 August 2011

This Hour costs 22 minutes

Another result I don't believe:
The amount of TV viewed in Australia in 2008 reduced life expectancy at birth by 1.8 years (95% uncertainty interval (UI): 8.4 days to 3.7 years) for men and 1.5 years (95% UI: 6.8 days to 3.1 years) for women. Compared with persons who watch no TV, those who spend a lifetime average of 6 h/day watching TV can expect to live 4.8 years (95% UI: 11 days to 10.4 years) less. On average, every single hour of TV viewed after the age of 25 reduces the viewer's life expectancy by 21.8 (95% UI: 0.3–44.7) min. This study is limited by the low precision with which the relationship between TV viewing time and mortality is currently known.
If people who watch six hours of daily television differ from other folks on margins other than TV-viewing, results here just might be overstated.

The paper just extrapolates to some actuarial tables the results of a prior article in Circulation. A couple things there to note:

  • There are no significant reported effects, after correcting for some health-related behaviours, for those watching less than four hours of television per day. And, saying that an hour costs 22 minutes is a bit nuts when the reference category is folks watching 2 hours of TV per day or less. Maybe there are effects for the "more than 6 hours per day" group. But that's hardly the same as saying that an hour of TV costs the moderate viewer 22 minutes.
  • The all-cause mortality relative risk even for folks in the >4 hour per day group has a confidence interval running from 1.04-2.05 after adjusting for those health behaviours that could be observed. That's almost touching the 1.0 mark; when the CI spans 1.0, results are insignificant. I wonder what would happen to those CIs on adjusting for more health behaviours.

HT: Alex Robson, who points to this journalistic account...