top of page

San Francisco Eviction Analysis

Nov 12, 2016
9 min read

The Motivation

With all the Silicon Valley housing market craze over the last few years, I was interested in something to do with San Francisco data since I’m now a denizen of the Bay Area. Kaggle originally had a dataset about crimes in Philly and I looked up to see what data was available to the public for San Francisco. While I found datasets for crime, transportation and other similar urban statistics, I was really interested in an eviction dataset given the dot-com boom, the Great Recession and the recent boom in movement to Silicon Valley and San Francisco, all of which I thought might have interesting impacts on eviction.

The Data

The dataset came as roughly 34,000 rows of data spanning from 1997 to present day with roughly 18 columns being a binary matrix of 1’s and 0’s indicating some reasons for why an eviction occurred. Each row also contained a neighborhood, location (in terms of longitude and latitude) and eviction date. While the data was relatively clean, there was definitely some work to be done with it in terms of combining columns, splitting factors, and aggregating counts.

I want to touch a bit more on the eviction reasons here. Some are pretty understandable and unsurprising such as:

  • Breach

  • Nuisance

  • Owner Move-In

Some others were things I might not have even thought of or revolved around evictions laws and policies such as:

  • Ellis Act Withdrawal: To quote the San Francisco Tenants Union, Ellis Act evictions generally are used to change the use of the building. Most Ellis evictions are used to convert rental units to condominiums, using loopholes in the condo law.

  • Good Samaritan Ends: Units can be rented to renters for very low cost by landlords when renters have been displaced as a result of some disaster.

First Impressions

I’m quickly learning that being handed a dataset with certain factors makes my mind run wild with possibilities of things I want to explore. I also started this project in Python before quickly switching over to R when I realized some of the work I wanted to do. It’s becoming apparent to me that while Python is faster iteratively and better for cleaning data, working with chunks of data, subsetting it, and plotting is really R’s strong suit.

The presence of a location factor immediately made me think about plotting data points on a map of San Francisco. I’d never worked with maps and plotting via R or Python before so the prospect was both a bit intimidating and exciting at the same time. Additionally, the presence of an eviction date column as well as neighborhood and the reason for evictions immediately made me ask the following questions:

  1. What are the top reasons for eviction?

  2. Which neighborhoods have the most evictions?

  3. Given the top x reasons for eviction in the dataset, what are the top y neighborhoods for those evictions?

  4. Given the top x neighborhoods for eviction in the data, what are the top x reasons for those evictions?

  5. How did evictions look in each neighborhood over the years?

  6. Did the location or type of evictions have some kind of pattern throughout the years 1997-2016?

Cleaning

Right off the bat, I did some basic data scrubbing. I dropped fields for “City”, “State”, and “Address” since this was all San Francisco data and I had the longitude/latitude anyways. I also split out the year from the eviction date for easier access later and split the location factor (which came as a longitude and latitude tuple) into separate columns.

I also created a separate dataframe containing the sum aggregate matrix of evictions by neighborhood since I knew I wanted to create some grouped bar charts later. With 41 unique neighborhoods and 18 unique eviction reasons, the data frame basically had rows for neighborhoods and columns were eviction reasons where each index (i, j) was the total number of evictions for neighborhood i for reason j.

Exploring the data

Time to take a peek and answer some questions. I decided to use a heat map to quickly glean what the top eviction reasons were and which were the top neighborhoods. While I expected certain reasons to be much more prevalent, I didn’t expect the skew that is shown below. In the heat map below, darker blues represent more evictions. As it turns out, a good majority of evictions were for just a few reasons. The spread of neighborhoods for an eviction reason was decently uniform though; at least more uniform than the spread of reasons.

Heat map of SF Neighborhoods and Evictions. Darker hues indicate greater rates.

Following up on that, I wanted to compare neighborhoods and eviction reasons in a slightly more quantitative depiction since heatmaps tend to be more qualitative. I wrote a set of functions that would return data frames that would allow me to directly answer questions 1-4 above. The graphs that emerged were quite interesting. San Francisco has a reputation for having neighborhoods that have very distinct and unique personalities as well as socioeconomic differences. Sure enough, when I plotted the top 10 neighborhoods for the top 6 eviction reasons for example, there wasn’t a consistent pattern for eviction reasons. Instead, certain neighborhoods had eviction reasons that dominated other reasons. To illustrate, while the top 4 reasons for eviction were Breach, Nuisance, Owner Move In and Ellis Act Withdrawal, Tender had a disproportionately large percentage of evictions as a result of Nuisance while Sunset/Parkside had an enormous amount Owner Move In relative to the total evictions in the neighborhood.

Capturing Neighborhood Characteristics

One aspect that I wanted to explore was how neighborhoods were affected by eviction throughout the years. With some of the aforementioned large events like the Great Recession, I expected evictions in most neighborhoods to drop as landlords would want to keep their tenants and development/remodeling to drop. I also expected that with any kind of job boom, that certain “hot” neighborhoods seen as desirable would see a boom in evictions. I thought, as touchy a subject as it might be, gentrification as a result the recent surge in Silicon Valley would lead to greater eviction rates as landlords might potentially search for reasons to evict tenants in search of new tenants who would pay a higher rent.

  • SF Evictions by Month
  • SF Evictions by Year

SF Evictions by Neighborhood from 1997-2016

The initial plot of evictions in San Francisco over time supported my hunch. The 1998-1999 period saw a large boom in evictions while the Great Recession saw a slump in evictions before another rise starting around 2010. Then, thanks to the multiplot code that was available online, I managed to squeeze plots for every single neighborhood into one giant plot. The plots generated were quite interesting: not every neighborhood followed the overall pattern that the city as a whole did. While it wasn’t entirely unsurprising, it was interesting nonetheless. As expected, relatively non-residential areas like Presidio, Golden Gate Park, Treasure Island and McLaren Park saw very few evictions over the years. Certain neighborhoods like Presidio Heights, Inner Richmond and Glen Park saw a general downwards trend while certain neighborhoods like Outer Mission, Mission, and Visitacion Valley followed the trajectory I expected: a spike around 2000, followed by a slump into the Great Recession before an increase again after roughly 2008. Once again, I want to just point out as of the data, the year 2016 was not over yet which may play a large part in the drop we're seeing in the plots above. Also worth noting is that the "Illegal Use" of a unit is no longer considered a just cause for eviction as of November 5th, 2015 according to the San Francisco Tenants Union which may also play a part in the dip in evictions for early 2016 data.

One thing I want to be wary of is assigning general trends in the economy as reasons for eviction. In fact, I need to dig more into the data to see if there is some lag to evictions during the Great Recession. I’m particularly curious to see if there was a bit of lag before Non-payment evictions might have potentially risen during the Great Recession as I would think that lack of funds for payment might not become limiting for tenants until a period of time into the Great Recession.

Where Are These Evictions

Lastly, I wanted to take a look at where these evictions were taking place. As I mentioned before, I had never worked with maps in either Python or R before but the ggmap package in R made it exceedingly simple. I ended up pulling up a Google Maps satellite roadmap of San Francisco and plotting the top 5 eviction reasons per year and outputting each plot before conglomerating them into a gif hoping that an animated picture would be helpful in illustrating the movement of eviction data points as well as potential clusters of similar eviction reasons.

  • San Francisco Evictions in 1997

Something worth noting here is that there were no failure to sign renewals during the 2004-2008 period. Another is the pocket of evictions that popped up the Parkmerced (just east of Lake Merced Park) area during 2010 and 2011. A quick Google search brought up articles about many Section 8 tenants being evicted during that time.

I think the most interesting trend really starts in 2010. Starting from 2010, the evictions in the northeast corner of San Francisco really starts to skyrocket. This is especially apparent near and around the intersection of Highway 101 and Market Street (to those unfamiliar, about 4th street northwest of highway 80). Coincidentally, Market Street is exactly where the BART runs. Given the boom of the technology sector and the proximity of Highway 101 and BART, it's no surprise that more evictions started popping up around that area. At the same time, certain neighborhoods have experienced decently consistent evictions over time or are characterized by a certain eviction type (i.e. The Tenderloin and nuisance evictions).

News articles reveal that San Franciscans definitely felt the pain of the rise in evictions. Many articles written around 2014/2015 reiterated the rise in evictions. Especially worth noting was the rise in Ellis Act evictions which, as the SF Examiner explained, was largely "due to real estate speculators who have just purchased a rent-controlled building." A quick slice of my dataset revealed that between 2008-2013, the number of Ellis Act Evictions progressed as follows: 48, 68, 54, 99, 231. A plethora of articles online have stories about landlords using loopholes in contracts or unfairly accusing tenants of nuisance or breach of contract in order to evict them.

What Didn't Work

I really wanted to do some kind of supervised learning algorithm here. For those new to analytics, supervised algorithms are used when you know the outcome of what you are trying to predict. In my case, I wanted to see if I could predict the type of eviction based on the covariates available to me.

Since this was a classification problem, I quickly settled on using trees to try and predict the eviction type. On first pass, the predictors that I thought would be most useful were Neighborhood, Latitude, Longitude and Year. As I quickly discovered, R's tree package as well as ctree in the 'party' package only support 32 levels for a factor variable (and I had about 50 neighborhoods or so). A workaround was to use a randomForest (an ensemble method commonly used to reduce variation with trees since those are prone to overfitting without pruning) which had a higher cap on the number of levels per factor variable.

What resulted was...disappointing at best. I tried a few permutations of predictor variables (using zip instead of latitude/longitude, removing year, etc). But nearly every model (using 500 trees) yielded an error rate around 50-60% which is completely unacceptable. It then occurred to me that the the types of eviction were extremely skewed to the top few eviction reasons. However, even after filtering the dataset to only the top 5 eviction reasons (the ones seen in the maps above), the error rate still only dropped to 44.29% when utilizing Neighborhood, Year, Latitude and Longitude as the covariates which is clearly not a very good model.

However, if we try to predict the Neighborhood using a random forest while using Reason, Year, Latitude and Longitude, the error rate plummets to just 1.32%. The very obvious issue here is that the Latitude and Longitude are extremely strong predictors. Despite the fact that strict lateral/longitudinal boundaries were not defined for the neighborhoods, they still provide extremely strong predictive power for ascertaining the neighborhood. I thought this would be an interesting case to point out since it exemplifies flawed methodology.

The Lessons

This was once again another project focused around exploratory data analysis and trying to illustrate and visualize data. It was interesting to see the eviction trends and how they support the general social and economic changes for San Francisco over the time period.

Unfortunately, trees and random forests did not provide a strong predictive model. Since choices for classification algorithms are relatively limited when there exists categorical data, I'll potentially need to look into transforming the data to produce a better model. A next good option would be using K nearest neighbors with the Hamming Distance for the categorical variables. However, this means that the numerical variables will need to be standardized to a [0,1] domain. Another possibility is exploring the k-prototypes clustering algorithm as identified in this paper by ZheXue Huang.

Next Steps

RStudio recently released v1.0 and I just wanted to give them credit for making such an incredible IDE. I also discovered more about R Markdown, Shiny and the knitr package. I’m digging into creating an interactive Shiny app to host on an Amazon instance. As always to any readers, please let me know of any suggestions, critiques and questions!

Comments


  • linkedin
  • generic-social-link

©2016 by Jason Wang. Proudly created with Wix.com

bottom of page