Student Spotlight: Racial Bias in the NYPD Stop-and-frisk Policy

Donald Trump recently came out in favor of an old New York Police Department’s (NYPD) “stop-and-frisk” policy that allowed police officers to stop, question and frisk individuals for weapons or illegal items. This policy was under harsh criticism for racial profiling and was declared in violation of the constitution by a federal judge in 2013.

An earlier post by Krit Petrachainan showed a potential racial discrimination against African-Americans within different precincts. Expanding on this topic, we decided to look at data in 2014, one year after the policy had been reformed, but when major official policy changes had not yet taken place.

More specifically, this study examined whether race (Black or White) influenced the chance of being frisked after being stopped in NYC in 2014 after taking physical appearance, population distribution among different race suspect, and suspected crime types into account.

2014 Data From NYPD and Study Results

For this study, we used the 2014 Stop, Question and Frisked dataset retrieved from the city of New York Police Department. After data cleaning, the final dataset has 22054 observations. To address our research question, we built a logistic regression model and ran a drop-in-deviance test to determine the importance of Race variable in our model.

Our results suggest that after the suspect is stopped, race does not significantly influence the chance of being frisked in NYC in 2014. A drop-in-deviance tests after creating a logistic regression model predicting the likelihood of being frisked gave a G-statistic of 8.99, and corresponding p-value of 0.061. This marginally significant p-value shows we do not have enough evidence to conclude that adding terms associated with Race improves the predictive power of the model.

Figure 1. Logistic regression plot predicting probability of being frisked from precinct population Black, compared across race

To better visualize the relationship in interactions between race and other variables, we created logistic regression plots predicting the probability of being frisked from either Black Pop or Age, and bar charts comparing proportion of suspects frisked across sex and race.

Interestingly, given that the suspects are stopped, as the precinct proportion of Blacks increases, both Black and White suspects are more likely to be frisked. Furthermore, this trend is more profound for Black than White suspects (Figure 1).

Additionally, young Black suspects are much more likely than their White counterparts to be frisked, given that they are stopped. This difference diminishes as suspect age increases (Figure 2).

Figure 2. Logistic regression plot predicting probability of being frisked from age, compared across age

Finally, male suspects are much more likely to be frisked than females, given they are stopped (Figure 3). However, the bar charts indicate that the effect of race on the probability of being frisked does not depend on gender.

Figure 3. Proportion frisked by race, compared across sex

Is stop-and-frisk prone to racial bias?

Our results suggest that given that the suspect is stopped, after taking other external factors into account, race does not significantly influence the chance of being frisked in NYC in 2014. However, after looking at relationships between race and precinct population Black, age, and sex, there is a possibility that the NYPD “stop-and-frisk” practices are prone to racism, posing threat to minority citizens in NYC. It is crucial that the NYPD continue to evaluate its “stop-and-frisk” policy and make appropriate changes to the policy and/or police officer training in order to prevent racial profiling at any level of investigation.

*** This study by Linh Pham, Takahiro Omura, and Han Trinh won 2nd place in the 2016 Undergraduate Class Project Competition (USCLAP).

5 Must-See TED Talks on Data Visualization!

Data visualization is crucial in understanding data and identifying hidden connections that matter. Below are 5 TED talks on data visualization you don’t want to miss!

1. Hans Rosling: The best stats you’ve ever seen

Han Rosling, cofounder of the Gapminder Foundation, developed the Trendalyzer software that converts international statistics – such as life expectancy and child mortality rate – into innovative, interactive graphics. The statistics guru is a strong advocate for public access to data and the development of tools that make it accessible and usable for all.  In this classic talk, Rosling highlights the importance of data in debunking myths about the gap between developed countries and the so-called “developing world.” Even though the talk was filmed 10 years ago, it still carries very important and relevant messages.

2. David McCandless: The beauty of data visualization

In this visually captivating talk, data journalist David McCandless suggests that data visualization is a quick solution to our current problem of information overload. Visualizations allow us to see the hidden patterns, identify connections that matter, and tell stories with data. To McCandless, “even when the information is terrible, the visual can be quite beautiful”; this is a controversial claim, however, since the main goal of data visualization should be to communicate information effectively through graphical means.

3. Dave Troy: Social maps that reveal a city’s intersections – and separations

A serial entrepreneur and data-viz fan, Dave Troy takes a people-focused approach to data visualization. Troy has been mapping tweets among city dwellers, revealing what connects communities and what separates them – above and beyond demographic factors such as race or ethnicity. He compares a city to a “giant high school cafeteria” and suggests that we see “how everybody arranged themselves in a seating chart”, arguing that “maybe it’s time to shake up the seating chart a little bit” to reshape our cities.

4. Eric Berlow & Sean Gourley: Mapping ideas worth spreading

An ecologist and a physicist, Eric Berlow and Sean Gourley, collaborate in this presentation to create stunning 3D visualizations demonstrating the interconnectedness of ideas. Taking 4,000 TEDx talks from 147 countries representing 50 languages, they explore their “meme-omes” – the mathematical structures that underlie the ideas behind these talks – and discover similarities between seemingly unconnected topics. Berlow and Gourley also broke down complex themes into multiple more specific ones, seeing what topics resonated with viewers and what kind of audience looked at what topic. To Gourley, mapping ideas in this way will help us “to see what’s being said, to see what’s not being said, and to be a little bit more human and, hopefully, a little smarter.”

5. Manuel Lima: A visual history of human knowledge

Founder of Manuel Lima, described by Wired Magazine as “the man who turns data into art,” explains the visual metaphor shift from the tree to the network as “a new lens to understand the world around us.” Lima argues that the tree – an important tool to map everything from genealogy to systems of law to Darwin’s “Tree of Life” – is being replaced by a new metaphor – the network. Rigid structures are evolving into interdependent systems, and networks emerge to embody the nonlinearity, decentralization, interconnectedness, and multiplicity of ideas and knowledge. The shift in visual metaphor also represents a new way of thinking – one that is critical for us to solve many complex problems we are facing.

How Traditional Introductory Statistics Textbooks Fail to Serve Social Science Undergraduates

When no weighting variable is used, the estimate is that about 50% of the population know the Jewish Sabbath starts on Friday.

No weighting variable: the estimate is that about 50% of the population knows that the Jewish Sabbath starts on Friday.

When the data is appropriately weighted, the estimate changes by about 5 percentage points.

Appropriately weighted data: The estimate changes by about 5 percentage points, suggesting that only 45% of the population knows the correct start time.

Full disclosure: I approach this topic simultaneously from the perspective of a social scientist and as the instructor of a traditional introductory statistics class for over twenty years. I am, thus, myself part of the problem. While I am mainly following the dictates of some of the most popular text books, it is fully within my power to diverge from the book. When I do not do so, it is really my own fault—a sheep following the sheep dogs.

Our worst failure as statistics teachers is to teach as if all or most of the data that our students will engage with in their future careers are from simple random samples.