<?xml version="1.0" encoding="utf-8"?>


<feed xmlns="http://www.w3.org/2005/Atom" >
  <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator>
  <link href="https://alexgude.com/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://alexgude.com/" rel="alternate" type="text/html" />
  <updated>2026-08-10T15:42:52-07:00</updated>
  <id>https://alexgude.com/feed.xml</id>

  
  
  

  
    <title type="html">Alex Gude</title>
  

  
    <subtitle>Technology, data science, machine learning, and more!</subtitle>
  

  
    <author>
        <name>Alexander Gude</name>
      
      
    </author>
  

  
  
  
  
  
  
    <entry>
      

      <title type="html">My Favorite Books of 2025</title>
      <link href="https://alexgude.com/blog/favorite-books-of-2025/" rel="alternate" type="text/html" title="My Favorite Books of 2025" />
      <published>2026-01-04T00:00:00-08:00</published>
      <updated>2026-01-04T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/favorite_books_of_2025</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/favorite-books-of-2025/"><![CDATA[<p>Last year I reviewed 45 books and 1 computer game. I mostly stuck to science
fiction, and I also re-read a few books, which was especially enjoyable
because I picked up so much more the second time through.</p>

<p>Following the tradition of <a href="/blog/favorite-books-of-2023/">2023</a> and <a href="/blog/favorite-books-of-2024/">2024</a>, here are my
favorite books from 2025:</p>

<h2 id="disco-elysium"><cite class="book-title">Disco Elysium</cite></h2>

<p>“<a href="/books/disco_elysium/"><cite class="book-title">Disco Elysium</cite></a> isn’t a book, Alex!” Ok, sure, but it <strong>is</strong> one of the literary
masterpieces of the 21st century. It does a miraculous job of immersing the
player in its world, exploring the grip the past has on the present, and
creating some of the deepest and most fully developed characters in any
medium. It’s my favorite game of all time, and one of my favorite literary
works, period.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/disco_elysium/">
      <img src="https://alexgude.com/books/covers/disco_elysium.jpg" alt="Book cover of Disco Elysium." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/disco_elysium/">
      <strong><cite class="book-title">Disco Elysium</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/robert_kurvitz/"><span class="author-name">Robert Kurvitz</span></a> <abbr class="etal">et al.</abbr></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="game-title">Disco Elysium</cite>, written by <span class="author-name">Robert Kurvitz</span> <abbr class="etal">et al.</abbr>, is a role-playing game
produced by ZA/UM. It’s the story of Harrier “Harry” Du Bois, a man who wakes
up with no memories and has to solve a murder while learning who he is.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-hyperion-cantos-by-dan-simmons">The <span class="book-series">Hyperion Cantos</span> by <span class="author-name">Dan Simmons</span></h2>

<p>I didn’t love the <a href="/books/series/hyperion_cantos/"><span class="book-series">Hyperion Cantos</span></a> the first time I read them. I found <a href="/books/the_fall_of_hyperion/"><cite class="book-title">The Fall of Hyperion</cite></a>,
the easier and more straightforward of the two, to be great, but I didn’t
really get <a href="/books/hyperion/"><cite class="book-title">Hyperion</cite></a>. This year I re-read them for my book club and focused on
finding and understanding the influences from <span class="author-name">John Keats</span>’s works. That
completely transformed the experience. I <strong>LOVED</strong> <a href="/books/hyperion/"><cite class="book-title">Hyperion</cite></a>, seeing it as far
deeper and more complex than I did the first time through.</p>

<p>This year I’ll be tackling <a href="/books/endymion/"><cite class="book-title">Endymion</cite></a> and <a href="/books/the_rise_of_endymion/"><cite class="book-title">The Rise of Endymion</cite></a>, which I’ve
heard don’t live up to the greatness of the first two, but I’m still hopeful.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/hyperion/">
      <img src="https://alexgude.com/books/covers/hyperion.jpg" alt="Book cover of Hyperion." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/hyperion/">
      <strong><cite class="book-title">Hyperion</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/dan_simmons/"><span class="author-name">Dan Simmons</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Hyperion</cite> is <span class="author-name">Dan Simmons</span>’s masterpiece. It is the first book in his <span class="book-series">Hyperion Cantos</span>. It follows seven pilgrims as
they travel to the Time Tombs on Hyperion to petition the Shrike. Along the
way, each tells their own story, weaving together history, myth, and prophecy
to tell of the impending downfall of man.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_fall_of_hyperion/">
      <img src="https://alexgude.com/books/covers/the_fall_of_hyperion.jpg" alt="Book cover of The Fall of Hyperion." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_fall_of_hyperion/">
      <strong><cite class="book-title">The Fall of Hyperion</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/dan_simmons/"><span class="author-name">Dan Simmons</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Fall of Hyperion</cite>, by <span class="author-name">Dan Simmons</span>,
is the second book in the <span class="book-series">Hyperion Cantos</span>, but really
it’s the second half of <a href="/books/hyperion/"><cite class="book-title">Hyperion</cite></a>. It brings the seven
pilgrims’ story to an end and depicts the war between the TechnoCore, the
Ousters, and the Hegemony.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-hydrogen-sonata-by-iain-m-banks"><cite class="book-title">The Hydrogen Sonata</cite> by <span class="author-name">Iain M. Banks</span></h2>

<p><a href="/books/authors/iain_m_banks/"><span class="author-name">Banks</span>’s</a> <a href="/books/series/culture/"><span class="book-series">Culture</span></a> series was one of my <a href="/blog/favorite-books-of-2024/">favorite reads in
2024</a>. I deliberately delayed finishing it, spacing out the final
few books because I knew that once I was done, it would be over for good. <a href="/books/the_hydrogen_sonata/"><cite class="book-title">The Hydrogen Sonata</cite></a> was a fitting end to the series. It restates many of the themes <a href="/books/authors/iain_m_banks/"><span class="author-name">Banks</span></a> first explored in <a href="/books/consider_phlebas/"><cite class="book-title">Consider Phlebas</cite></a>, but with a more hopeful bent.</p>

<p>It also feels like a fitting final book for <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a> himself, with its
message that life only has the meaning you give it.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_hydrogen_sonata/">
      <img src="https://alexgude.com/books/covers/the_hydrogen_sonata.jpg" alt="Book cover of The Hydrogen Sonata." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_hydrogen_sonata/">
      <strong><cite class="book-title">The Hydrogen Sonata</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Hydrogen Sonata</cite>, by <span class="author-name">Iain M. Banks</span>,
is the tenth and final book in the <span class="book-series">Culture</span> series. It
explores the last days of the Glitz people as they prepare to Sublime.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-teixcalaan-series-by-arkady-martine">The <span class="book-series">Teixcalaan</span> series by <span class="author-name">Arkady Martine</span></h2>

<p>Another pair of books picked by my book club. I really enjoyed <a href="/books/authors/arkady_martine/"><span class="author-name">Martine</span>’s</a>
subtle worldbuilding and the deep integration of themes and motifs across <a href="/books/a_memory_called_empire/"><cite class="book-title">A Memory Called Empire</cite></a> and <a href="/books/a_desolation_called_peace/"><cite class="book-title">A Desolation Called Peace</cite></a>. The writing itself is beautiful, and often reminded
me of <a href="/books/authors/gene_wolfe/"><span class="author-name">Gene Wolfe</span>’s</a> <a href="/books/series/the_book_of_the_new_sun/"><span class="book-series">The Book of the New Sun</span></a>, with its mix of archaic prose and intentional
vagueness.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/a_memory_called_empire/">
      <img src="https://alexgude.com/books/covers/a_memory_called_empire.jpg" alt="Book cover of A Memory Called Empire." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/a_memory_called_empire/">
      <strong><cite class="book-title">A Memory Called Empire</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/arkady_martine/"><span class="author-name">Arkady Martine</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">A Memory Called Empire</cite>, by <span class="author-name">Arkady Martine</span>,
is the first book in the <span class="book-series">Teixcalaan</span> series. It follows
Mahit Dzmare, an ambassador from the space station Lsel, as she tries to save
her home from being annexed by the Teixcalaanli empire.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/a_desolation_called_peace/">
      <img src="https://alexgude.com/books/covers/a_desolation_called_peace.jpg" alt="Book cover of A Desolation Called Peace." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/a_desolation_called_peace/">
      <strong><cite class="book-title">A Desolation Called Peace</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/arkady_martine/"><span class="author-name">Arkady Martine</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">A Desolation Called Peace</cite>, by <span class="author-name">Arkady Martine</span>,
is the second book in the <span class="book-series">Teixcalaan</span> series. It tells the
story of Mahit and Three Seagrass trying to stop the war between the
Teixcalaanli Empire and a mysterious alien race.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="roadside-picnic-by-arkady-and-boris-strugatsky"><cite class="book-title">Roadside Picnic</cite> by <span class="author-name">Arkady</span> and <span class="author-name">Boris Strugatsky</span></h2>

<p>I picked up <a href="/books/authors/arkady_strugatsky/"><span class="author-name">Arkady</span></a> and <a href="/books/authors/boris_strugatsky/"><span class="author-name">Boris Strugatsky</span>’s</a> best-known novel, <a href="/books/roadside_picnic/"><cite class="book-title">Roadside Picnic</cite></a>, because I
wanted to broaden what I was reading. I ended up loving it. The characters
feel incredibly real, the dialogue has a simple but earnest quality, and the
atmosphere is completely different from Western sci-fi.</p>

<p>It wasn’t the only book of theirs that I read, but it was easily my favorite.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/roadside_picnic/">
      <img src="https://alexgude.com/books/covers/roadside_picnic.jpg" alt="Book cover of Roadside Picnic." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/roadside_picnic/">
      <strong><cite class="book-title">Roadside Picnic</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/arkady_strugatsky/"><span class="author-name">Arkady Strugatsky</span></a> and <a href="/books/authors/boris_strugatsky/"><span class="author-name">Boris Strugatsky</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Roadside Picnic</cite>, by brothers <span class="author-name">Arkady</span> and <span class="author-name">Boris Strugatsky</span>, is a Soviet sci-fi novel. It’s essentially four short
stories—each presented as a chapter—about the life of Redrick “Red”
Schuhart, a “stalker” who illegally enters an alien-contaminated Zone to
retrieve items for the black market.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-murderbot-diaries-by-martha-wells"><span class="book-series">The Murderbot Diaries</span> by <span class="author-name">Martha Wells</span></h2>

<p>Everyone loves <a href="/books/series/the_murderbot_diaries/"><span class="book-series">The Murderbot Diaries</span></a>, and I’m no exception. <a href="/books/authors/martha_wells/"><span class="author-name">Wells</span></a> hits the
perfect mix of pulpy action, fast pacing, and memorable characters, while
still digging into deep philosophical questions about what it means to be a
person. I’ve got a few more to go this year, and I’m really looking forward to
them. I fully expect to see them on my 2026 list.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/artificial_condition/">
      <img src="https://alexgude.com/books/covers/artificial_condition.jpg" alt="Book cover of Artificial Condition." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/artificial_condition/">
      <strong><cite class="book-title">Artificial Condition</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/martha_wells/"><span class="author-name">Martha Wells</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Artificial Condition</cite>, by <span class="author-name">Martha Wells</span>,
is the second book in <span class="book-series">The Murderbot Diaries</span>. It follows
Murderbot as it digs into its past and, once again, saves some scientists.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/rogue_protocol/">
      <img src="https://alexgude.com/books/covers/rogue_protocol.jpg" alt="Book cover of Rogue Protocol." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/rogue_protocol/">
      <strong><cite class="book-title">Rogue Protocol</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/martha_wells/"><span class="author-name">Martha Wells</span></a></span>
    <div class="book-rating star-rating-3" role="img" aria-label="Rating: 3 out of 5 stars. Good — I liked it" title="Good — I liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Rogue Protocol</cite>, by <span class="author-name">Martha Wells</span>,
is the third book in <span class="book-series">The Murderbot Diaries</span>. It follows
Murderbot as it investigates a GrayCris terraforming station and, you guessed
it, ends up saving a group of humans.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/exit_strategy/">
      <img src="https://alexgude.com/books/covers/exit_strategy.jpg" alt="Book cover of Exit Strategy." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/exit_strategy/">
      <strong><cite class="book-title">Exit Strategy</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/martha_wells/"><span class="author-name">Martha Wells</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Exit Strategy</cite>, by <span class="author-name">Martha Wells</span>,
is the fourth book in <span class="book-series">The Murderbot Diaries</span>. It wraps up
the GrayCris storyline as Murderbot returns to save its friends.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/network_effect/">
      <img src="https://alexgude.com/books/covers/network_effect.jpg" alt="Book cover of Network Effect." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/network_effect/">
      <strong><cite class="book-title">Network Effect</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/martha_wells/"><span class="author-name">Martha Wells</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Network Effect</cite>, by <span class="author-name">Martha Wells</span>,
is the fifth book in <span class="book-series">The Murderbot Diaries</span>. It’s the first
full-length novel in the series and features Murderbot getting kidnapped by
ART to rescue its crew.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="there-is-no-antimemetics-division-by-qntm"><cite class="book-title">There Is No Antimemetics Division</cite> by <span class="author-name">qntm</span></h2>

<p>I read <a href="/books/authors/qntm/"><span class="author-name">qntm</span>’s</a> <a href="/books/there_is_no_antimemetics_division_original/"><cite class="book-title">original edition</cite></a> edition in 2023 and thought it was packed
full of amazing ideas, but ultimately let down by a poorly written second act.
His rewrite this year fixed every single problem. The result is a tight,
cohesive book that brings its wild concepts to life. Highly recommend.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/there_is_no_antimemetics_division/">
      <img src="https://alexgude.com/books/covers/there_is_no_antimemetics_division.jpg" alt="Book cover of There Is No Antimemetics Division." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/there_is_no_antimemetics_division/">
      <strong><cite class="book-title">There Is No Antimemetics Division</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/qntm/"><span class="author-name">qntm</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">There Is No Antimemetics Division</cite>, by <span class="author-name">qntm</span>,
is a book about researchers trying to control dangerous antimemes—ideas that
can’t be thought—and how you might combat a foe you can’t even remember
exists.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="a-mote-in-shadow-by-a-n-alex"><cite class="book-title">A Mote in Shadow</cite> by <span class="author-name">A. N. Alex</span></h2>

<p>I found this book when my corner of BlueSky started talking about it, and I’m
glad I paid attention. <a href="/books/authors/a_n_alex/"><span class="author-name">Alex</span>’s</a> debut novel is a fantastic mix of hard
sci-fi and techno-thriller, set in a fractured human civilization in the
not-too-distant future. He’s working on the sequel now, and I’m very much
looking forward to it.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/a_mote_in_shadow/">
      <img src="https://alexgude.com/books/covers/a_mote_in_shadow.jpg" alt="Book cover of A Mote in Shadow." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/a_mote_in_shadow/">
      <strong><cite class="book-title">A Mote in Shadow</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/a_n_alex/"><span class="author-name">A. N. Alex</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">A Mote in Shadow</cite> is <span class="author-name">A. N. Alex</span>’s debut novel. It follows two down-on-their-luck outsiders dragged
into a war between shadowy mercenary groups: exobiologist Chaeyoung No, whose
disagreement with the scientific establishment leaves her in no position to
question a too-good-to-be-true offer to fund her research expedition; and
space hauler Frederik Obialo, who is more than willing to take a dangerous job
if it brings him closer to his dream of giving his daughter a permanent home.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-triumphant-by-linda-evans-robert-r-hollingsworth-and-david-weber"><cite class="book-title">The Triumphant</cite> by <span class="author-name">Linda Evans</span>, <span class="author-name">Robert R. Hollingsworth</span>, and <span class="author-name">David Weber</span></h2>

<p>I started reading the <a href="/books/series/bolo/"><span class="book-series">Bolo</span></a> books when I was a teen and absolutely loved
them. They feature giant tanks blowing stuff up, but, like <a href="/books/series/the_murderbot_diaries/"><span class="book-series">The Murderbot Diaries</span></a>,
they also broach deeper questions about the meaning of being alive and whether
honor and duty translate to machines. I recently started re-reading the series
for nostalgia, and while most of the books are serviceable but not great, <a href="/books/the_triumphant/"><cite class="book-title">The Triumphant</cite></a> is the clear exception.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_triumphant/">
      <img src="https://alexgude.com/books/covers/bolos_book_3_the_triumphant_1st_edition.jpg" alt="Book cover of The Triumphant." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_triumphant/">
      <strong><cite class="book-title">The Triumphant</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/linda_evans/"><span class="author-name">Linda Evans</span></a>, <a href="/books/authors/robert_r_hollingsworth/"><span class="author-name">Robert R. Hollingsworth</span></a>, and <a href="/books/authors/david_weber/"><span class="author-name">David Weber</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Triumphant</cite> is the twelfth book in the <span class="book-series">Bolo</span> series. It’s an anthology of Bolo stories written by three different
authors. They explore the emotional bond between a Bolo and the people around
them, and the dangers of caring too much about a machine built for war.</p>
    </div>
  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="book-reviews" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[I read 45 books in 2025, diving deep into sci-fi classics and modern hits. From the Hyperion Cantos and the final Culture novel to the philosophical action of Murderbot, plus one literary masterpiece that is actually a video game, here are my absolute favorites of the year.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/book-reviews/books_from_the_marburg_image_archive.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/book-reviews/books_from_the_marburg_image_archive.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">A Letter to my Alma Mater</title>
      <link href="https://alexgude.com/blog/a-letter-to-my-alma-mater/" rel="alternate" type="text/html" title="A Letter to my Alma Mater" />
      <published>2025-09-12T00:00:00-07:00</published>
      <updated>2025-09-12T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/a_letter_to_my_alma_mater</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/a-letter-to-my-alma-mater/"><![CDATA[<p>It was with dismay this morning that I read the Daily Cal article: <a href="https://www.dailycal.org/news/campus/uc-berkeley-turns-over-personal-information-of-more-than-150-students-and-staff-to-federal/article_a4aad3e1-bbba-42cc-92d7-a7964d9641c5.html"><cite class="newspaper">UC Berkeley turns over personal information of more than 150
students and staff to federal government</cite></a>. That you, the
inheritors of this university’s mantle, would willingly comply with a
political witch hunt reveals a profound and stunning ignorance of the
historical moment we are in.</p>

<p>Your failure of courage raises some fundamental questions. What is the point
of studying history if we learn nothing from it? What is the point of our own
history if we simply repeat the same mistakes we made in the past?</p>

<p>We tell ourselves myths about Mario Savio and the Free Speech Movement, about
the little circle of freedom in front of Sproul Hall, about how that history
determined who we are now. But it is clear now that you have chosen to betray
that legacy, siding not with Savio, but with the machine he warned us about.
You aren’t willing to put your careers, never mind your bodies, upon the
wheels and levers to make it stop.</p>

<p>You believe you’re saving the University. But you are destroying its soul.</p>

<p><strong>Alexander Gude</strong> (‘08)</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[In a stunning failure of courage, UC Berkeley has turned over the personal information of students and staff to the federal government. This decision is a profound betrayal of the university's legacy, particularly the Free Speech Movement. This is my letter to the administration.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/berkeley/berkeley_wheeler.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/berkeley/berkeley_wheeler.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">LLMs Make Python Scripts Free</title>
      <link href="https://alexgude.com/blog/llms-make-python-free/" rel="alternate" type="text/html" title="LLMs Make Python Scripts Free" />
      <published>2025-08-30T00:00:00-07:00</published>
      <updated>2025-08-30T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/llms_make_python_free</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/llms-make-python-free/"><![CDATA[<p>I first learned Python in 2005. My mentor at the time, Nao Suzuki, was
training me to do cosmology research and decided it would be a better use of
my time to learn Python instead of IDL. He also told me something that changed
how I thought about computers:</p>

<blockquote>
  <p>A computer’s job is to do work for you.</p>
</blockquote>

<p>I hadn’t realized that before. Until I learned to program, computers could
only do a small set of things for me—mainly the things other people had
already decided they should do. But when I learned to code, I suddenly had the
ability to make them do what I wanted.</p>

<p>Within a year I’d learned Python, picked up Vim, and switched my computers to
Ubuntu. I started writing little scripts to make my life easier, like backing
up my email, syncing files between my laptop and desktop, or converting
Wikipedia pages into an archival format.</p>

<p>Back then, each script took me hours to write. I had to find the right
libraries, learn their APIs, or sometimes write code from scratch. Because
they took so much work I saved every one, sometimes even sharing them with the
opensource community, just in case I ever needed them again. But all of that
changed in the last two years when I <a href="/blog/how-i-write-code-with-llms/">learned to use an LLM for
coding</a>.</p>

<p>Now a Python script takes 30 seconds. Maybe a minute if I want to read it over
and give some feedback. Gemini 2.5 Pro <em>almost always</em> writes a 100-line
script correctly on the first try, just from a short description of what I
want. It’s so fast and easy that I’ve stopped saving them—if I need another
one, I can just prompt for it again in half a minute.</p>

<p><strong>Python scripts are now free!</strong><sup style="anchor-name:--fnref-also" id="fnref:also"><a href="#fn:also" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> But I don’t think everyone has
realized this yet. Once again, <a href="/books/authors/william_gibson/"><span class="author-name">William Gibson</span></a> was right:</p>

<blockquote>
  <p>The future is already here – it’s just not very evenly distributed.</p>
</blockquote>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:also">

      <p>And not just Python, but any 300-line piece of code! I used Gemini to
completely re-write the backend of this blog in pure Ruby, cutting build
times from minutes to seconds, adding hundreds of tests to define and
enforce behavior, and I don’t even know Ruby! <a href="#fnref:also" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
          <category term="python" />
        
      

      

      
      
        <summary type="html"><![CDATA[I used to spend hours writing and saving little Python scripts. Now, with LLMs, I can get the same code in under a minute. Python scripts, and small programs in general, are effectively free.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/llm-free-python/a_robot_sells_pythons_of_many_colors.png" />
        <media:content medium="image" url="https://alexgude.com/files/llm-free-python/a_robot_sells_pythons_of_many_colors.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">How I Code With LLMs</title>
      <link href="https://alexgude.com/blog/how-i-write-code-with-llms/" rel="alternate" type="text/html" title="How I Code With LLMs" />
      <published>2025-05-31T00:00:00-07:00</published>
      <updated>2025-05-31T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/how_i_write_code_with_llms</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/how-i-write-code-with-llms/"><![CDATA[<p>I’ve been writing code professionally for 20 years. I started with Java, moved
on to Python, then C++, Scala, and back to Python, with a smattering of shell
scripting, PHP, and Rust in between. I’ve written code with a pen and paper,
Vim, and a full IDE. Now, I’m exploring writing code with <a href="https://en.wikipedia.org/wiki/Large_language_model">large language
models (LLMs)</a>, similar to <a href="/blog/how-i-write-with-llms-revised/">how I use them for writing</a>.</p>

<p>At work, I use <a href="https://block.github.io/goose/">Codename Goose</a>, which can directly interact with my
local machine and files, switch between multiple LLM APIs, and connect to our
infrastructure with <a href="https://en.wikipedia.org/wiki/Model_Context_Protocol">Model Context Protocol (MCP)</a> servers. At home, I’m
more limited, but use GPT-4o and Gemini 2.5 Pro to generate code (that I have
to copy back and forth between my browser and terminal).</p>

<p>After intentionally trying to shoehorn LLMs into my workflow to see where and
how they work, this is what I’ve learned.</p>

<h2 id="system-prompt">System Prompt</h2>

<p>The system prompt, which the model “reads” before it gets started, is how you
customize the way the model “thinks” and responds to you. I find a few things
are useful to include:</p>

<ol>
  <li><strong>A general overview of the task:</strong> This gives the LLM context about where
you’re headed, even as it works line by line and file by file. “We’re
refactoring a Jekyll website…”, “We’re building a set of machine learning
signals in Java 17…”, etc.</li>
  <li><strong>A layout of the project:</strong> Locations of key files, what goes where, etc.
Things like “a.py is where the main code lives, b.py provides helper
functions…”</li>
  <li><strong>A style guide:</strong> Telling the LLM how you prefer to use the programming
language. “Never use ternary operators”, “Always use <code class="language-plaintext highlighter-rouge">{}</code> with if
statements”, “Prefer functional solutions”, etc.</li>
</ol>

<p>A trick I use is, at the end of a session, I give the LLM the current system
prompt and ask it to update it—adding anything new it learned and fixing
anything that’s now out of date. Then I use that modified prompt when I spin
up a new LLM to work on the project.</p>

<h2 id="unit-tests">Unit Tests</h2>

<p>English isn’t a very precise way to define what you want code to do, but
that’s usually where I start, maybe with a few example code snippets. After
that, I ask the model to write <a href="https://en.wikipedia.org/wiki/Unit_testing">unit tests</a> for the code we just wrote.
This is where I make sure we’re on the same page.</p>

<p>I read through the unit tests and look for anything that doesn’t match what I
was expecting. Sometimes it’s easy to just change the test to match the
behavior I want. Other times, I try to understand why the LLM made the choice
it did, then discuss the decision with it to come to an agreement. I often
catch disagreements over when to throw errors, how to handle invalid inputs,
or other rare edge cases in this process.</p>

<p>Once the behavior is codified in the tests, I ask the model to fix the
original code so the tests pass. This iterative process of generation and
verification is <a href="/blog/good-uses-for-large-language-models/">where LLMs truly shine</a>. In this way, it’s
sort of an inverse of <a href="https://en.wikipedia.org/wiki/Test-driven_development">test-driven development</a>—I have the computer
write the code and then the tests. I find that having some code to start with
helps the LLM write broader, more useful tests, a practice whose importance
I’ve <a href="/blog/software-testing-for-data-science/">previously discussed in the context of data science</a>.</p>

<p>In some cases, I’ll even throw away the code, restart the LLM, and give it
just the tests as the spec for what it should write. This is particularly
useful when the model gets stuck and either won’t make a large enough change
to the code, or loops back and forth between essentially the same few
versions.</p>

<h2 id="iterate-and-advise">Iterate And Advise</h2>

<p>When I’m working with an LLM, I see my job more as a manager or mentor than an
engineer. My role is to define what we’re doing, then critique the result to
get it into shape. It’s a lot like working with a junior engineer or
intern—except one that can make a revision in 10 seconds instead of 4 hours.
This lets me iterate a lot to get the code working exactly how I want.</p>

<p>The things I focus on are mostly high-level, like data structures and
algorithms. The LLM often doesn’t need a lot of direction—just saying
“Couldn’t we replace the double for loop with a single loop and a hash table?”
is usually enough to get it on the right path. Only at the end, once the
structure is solid, do I focus on nitpicky issues or edit the text directly.</p>

<h2 id="the-future">The Future</h2>

<p>I’m still learning how to use LLMs for coding. The technology and tooling is
evolving so fast that I’m sure what I’m doing today will look antiquated in
six months. Still, I hope it gives others some ideas about how they can
use this technology to speed up their development.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[Over the past few years, I've experimented with coding alongside large language models. This post shares how I integrate them into my workflow.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/coding-llm/water_color_of_a_robot_writing_code.png" />
        <media:content medium="image" url="https://alexgude.com/files/coding-llm/water_color_of_a_robot_writing_code.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">How I Write with Large Language Models</title>
      <link href="https://alexgude.com/blog/how-i-write-with-llms-revised/" rel="alternate" type="text/html" title="How I Write with Large Language Models" />
      <published>2025-02-02T00:00:00-08:00</published>
      <updated>2025-02-02T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/how_i_write_with_llms_revised</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/how-i-write-with-llms-revised/"><![CDATA[<p><em>This is the final version of the post, refined through the editing process I
describe below. You can see my starting point by <a href="/blog/how-i-write-with-llms-revised/first-draft/">reading my first
draft</a> and comparing it with the <a href="/blog/how-i-write-with-llms-revised/llm-draft/">LLM-edited draft</a>.</em></p>

<p>ChatGPT 3.5 came out just over two years ago and sparked an explosion in Large
Language Model (LLM) development. Dozens of companies released their own
models, and the state of the art advanced by the hour.</p>

<p>At the time, <a href="/blog/how-i-write-with-chatgpt/">I wrote about how I used ChatGPT to write</a>. My
method was primitive. With years of experience and improvements in the models,
I have refined how I use LLMs to edit text. Here is my new method.</p>

<h2 id="drafting">Drafting</h2>

<p>I write the first draft entirely by hand. This helps me preserve my voice and
prevents the writing from being overly influenced by the LLM. It also supports
my main goal of writing: to clarify my thinking.</p>

<p>I used to write, edit, write, edit, and so on until I was nearly 100% happy
with my work, but now I stop earlier and let the LLM handle an editing pass.
Using the LLM early saves me multiple rounds of edits because they’ve become
so good at fixing spelling and grammatical errors and slightly tweaking my
writing without overpowering it.</p>

<h2 id="the-prompt">The Prompt</h2>

<p>The prompt you use with the LLM is important, because it strongly shapes how
the model edits your writing. The prompt keeps the machine from filling my
writing with phrases like <a href="https://arstechnica.com/ai/2024/07/the-telltale-words-that-could-identify-generative-ai-text/">“delve”, “showcasing”, and “underscores”</a>. I
currently use a slight variation of this prompt:</p>

<blockquote>
  <p>Help me edit this blog post I’m writing. Fix errors, make it clearer. Reword
to make the arguments and sentences more coherent. Use the same sort of
words I’m using, don’t substitute fancy synonyms. Maintain my voice. My work
is below. Keep the formatting and wrap your output in ```.</p>
</blockquote>

<h2 id="editing">Editing</h2>

<p>Once I have the LLM’s version, I put it side-by-side with my draft to compare.
Sometimes the edited version is perfect, and I’ll take an entire paragraph as
is. Other times, I’ll borrow ideas about how to structure a paragraph or a
transition, but I rewrite it in my own words. The model sometimes tries to fix
sections that are fine, and I simply ignore it.</p>

<p>After that, I go through another human editing pass to ensure my voice comes
through in every sentence. Sometimes, I will have the LLM focus on specific
sentences or paragraphs that still need work and iterate. Once I’m happy, I
commit my changes and publish.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[OpenAI's ChatGPT 3.5 transformed my writing process when it came out. After years of experience using it, I've further refined my method of using LLMs. This post explains how.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/chatgpt/202502001-robot.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/chatgpt/202502001-robot.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Favorite Books of 2024</title>
      <link href="https://alexgude.com/blog/favorite-books-of-2024/" rel="alternate" type="text/html" title="My Favorite Books of 2024" />
      <published>2025-01-01T00:00:00-08:00</published>
      <updated>2025-01-01T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/favorite_books_of_2024</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/favorite-books-of-2024/"><![CDATA[<p>In 2024, I finished 46 books and joined a sci-fi book club. The club has been
a fantastic way to explore books outside my usual reading habits and has
encouraged me to think more critically about writing, thematic connections,
and the author’s intent—so I have plenty to share during discussions.
Following the tradition of <a href="/blog/favorite-books-of-2023/#blindsight-by-peter-watts">last year’s post</a>, here are my favorite
reads from 2024:</p>

<h2 id="echopraxia-by-peter-watts"><cite class="book-title">Echopraxia</cite> by <span class="author-name">Peter Watts</span></h2>

<p><a href="/books/echopraxia/"><cite class="book-title">Echopraxia</cite></a> is the sequel to <a href="/books/blindsight/"><cite class="book-title">Blindsight</cite></a>, one of my <a href="/blog/favorite-books-of-2023/#blindsight-by-peter-watts">favorite books from last
year</a>. Although it wasn’t as well received by readers as <a href="/books/blindsight/"><cite class="book-title">Blindsight</cite></a>,
it was my favorite read of the year. In fact, I liked it <em>more</em> than <a href="/books/blindsight/"><cite class="book-title">Blindsight</cite></a>
because of its complex storyline, intriguing characters, and the broader
perspective it offers on the world.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/echopraxia/">
      <img src="https://alexgude.com/books/covers/echopraxia.jpg" alt="Book cover of Echopraxia." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/echopraxia/">
      <strong><cite class="book-title">Echopraxia</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_watts/"><span class="author-name">Peter Watts</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Echopraxia</cite>, by <span class="author-name">Peter Watts</span>,
is the second book in the <span class="book-series">Firefall</span> series, unfolding at
roughly the same time as <a href="/books/blindsight/"><cite class="book-title">Blindsight</cite></a>. It follows
parasitologist Daniel Brüks, who gets unwillingly dragged into a conflict
between multiple transhuman factions, travels to the <em>Icarus</em> station orbiting
the sun, and eventually back to Earth.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-culture-series-by-iain-m-banks">The <span class="book-series">Culture</span> Series by <span class="author-name">Iain M. Banks</span></h2>

<p>I <strong>hated</strong> <a href="/books/consider_phlebas/"><cite class="book-title">Consider Phlebas</cite></a> when I read it in 2023, but I gave the <a href="/books/series/culture/"><span class="book-series">Culture</span></a>
series another chance because I had spent the better part of two decades
waiting to get my hands on the books. I’m glad I did, because I read <a href="/books/the_player_of_games/"><cite class="book-title">The Player of Games</cite></a>
with my book club and loved it! My favorite from the series was <a href="/books/surface_detail/"><cite class="book-title">Surface Detail</cite></a>,
closely followed by <a href="/books/use_of_weapons/"><cite class="book-title">Use of Weapons</cite></a>. I even enjoyed <a href="/books/inversions/"><cite class="book-title">Inversions</cite></a>, one of the least
well-reviewed of <a href="/books/authors/iain_m_banks/"><span class="author-name">Banks</span>’s</a> books. It will be bittersweet to finish the
series with <a href="/books/the_hydrogen_sonata/"><cite class="book-title">The Hydrogen Sonata</cite></a> in 2025.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_player_of_games/">
      <img src="https://alexgude.com/books/covers/the_player_of_games.jpg" alt="Book cover of The Player of Games." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_player_of_games/">
      <strong><cite class="book-title">The Player of Games</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Player of Games</cite>, by <span class="author-name">Iain M. Banks</span>’s, is the second novel in the <span class="book-series">Culture</span> series. It tells the story of Jernau Morat Gurgeh, a master game player who is
recruited to play Azad, an incredibly complex game that serves as the basis
for the Empire of Azad’s entire government.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/inversions/">
      <img src="https://alexgude.com/books/covers/inversions.jpg" alt="Book cover of Inversions." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/inversions/">
      <strong><cite class="book-title">Inversions</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Inversions</cite>, by <span class="author-name">Iain M. Banks</span>,
is the sixth book in the <span class="book-series">Culture</span> series, but it is very
different from typical Culture novels: there are no spaceships and almost no
advanced technology. Instead, it follows Culture citizens DeWar and Vosill as
they manipulate a medieval society.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/look_to_windward/">
      <img src="https://alexgude.com/books/covers/look_to_windward.jpg" alt="Book cover of Look to Windward." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/look_to_windward/">
      <strong><cite class="book-title">Look to Windward</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Look to Windward</cite>, by <span class="author-name">Iain M. Banks</span>,
is the seventh book in the <span class="book-series">Culture</span> series. It explores
the aftermath of the Idiran–Culture War and Chelgrian civil war.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/surface_detail/">
      <img src="https://alexgude.com/books/covers/surface_detail.jpg" alt="Book cover of Surface Detail." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/surface_detail/">
      <strong><cite class="book-title">Surface Detail</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/iain_m_banks/"><span class="author-name">Iain M. Banks</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Surface Detail</cite>, by <span class="author-name">Iain M. Banks</span>,
is the ninth book in the <span class="book-series">Culture</span> series. It follows
Lededje Y’breq as she seeks revenge for her own murder, set against the
backdrop of a galactic conflict over virtual hells.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="a-fire-upon-the-deep-by-vernor-vinge"><cite class="book-title">A Fire Upon The Deep</cite> by <span class="author-name">Vernor Vinge</span></h2>

<p><a href="/books/a_fire_upon_the_deep/"><cite class="book-title">A Fire Upon The Deep</cite></a> is a nostalgic
favorite that I first read about 20 years ago and reread this year for my book
club. It does a fantastic job of telling a story that feels small and personal
while having galaxy-spanning implications. The Zones of Thought concept is
also a unique way to structure the galaxy and explore how it shapes
civilizations.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/a_fire_upon_the_deep/">
      <img src="https://alexgude.com/books/covers/a_fire_upon_the_deep.jpg" alt="Book cover of A Fire Upon The Deep." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/a_fire_upon_the_deep/">
      <strong><cite class="book-title">A Fire Upon The Deep</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/vernor_vinge/"><span class="author-name">Vernor Vinge</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">A Fire Upon The Deep</cite> is a sci-fi novel by <span class="author-name">Vernor Vinge</span>. It tells the story of the Blight—a
galactic-scale, transcendent evil—and the humans racing to stop it.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-cheela-series-by-robert-l-forward">The <span class="book-series">Cheela</span> Series by <span class="author-name">Robert L. Forward</span></h2>

<p>The <a href="/books/series/cheela/"><span class="book-series">Cheela</span></a> series consists of two hard sci-fi novels by <a href="/books/authors/robert_l_forward/"><span class="author-name">Forward</span></a>:
<a href="/books/dragons_egg/"><cite class="book-title">Dragon’s Egg</cite></a> and <a href="/books/starquake/"><cite class="book-title">Starquake</cite></a>. Even though the Cheela are <em>extremely</em> alien—living
on the surface of a neutron star and experiencing time a million times faster
than humans—their characters still pulled me in. It was exciting to watch
them build their civilization from hunter-gatherers to a spacefaring society.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/dragons_egg/">
      <img src="https://alexgude.com/books/covers/dragons_egg.jpg" alt="Book cover of Dragon&#39;s Egg." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/dragons_egg/">
      <strong><cite class="book-title">Dragon’s Egg</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/robert_l_forward/"><span class="author-name">Robert L. Forward</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Dragon’s Egg</cite> is a hard sci-fi novel by <span class="author-name">Robert L. Forward</span>. It is the story of first
contact between humans and the Cheela: beings who live on a neutron star.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/starquake/">
      <img src="https://alexgude.com/books/covers/starquake.jpg" alt="Book cover of Starquake." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/starquake/">
      <strong><cite class="book-title">Starquake</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/robert_l_forward/"><span class="author-name">Robert L. Forward</span></a></span>
    <div class="book-rating star-rating-3" role="img" aria-label="Rating: 3 out of 5 stars. Good — I liked it" title="Good — I liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Starquake</cite> is the second book in the <span class="book-series">Cheela</span> series by <span class="author-name">Robert L. Forward</span>. It follows
the Cheela as they rescue the humans and rebuild after a devastating
starquake.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="childhoods-end-by-arthur-c-clarke"><cite class="book-title">Childhood’s End</cite> by <span class="author-name">Arthur C. Clarke</span></h2>

<p>Another nostalgic read, <a href="/books/authors/arthur_c_clarke/"><span class="author-name">Clarke</span>’s</a> <a href="/books/childhoods_end/"><cite class="book-title">Childhood’s End</cite></a> was one of the first sci-fi
books I ever read. Revisiting it years later was a treat, and I’m happy to say
it holds up. The focus on humans’ psychic abilities feels a little dated, but
<a href="/books/authors/arthur_c_clarke/"><span class="author-name">Clarke</span>’s</a> crisp writing kept me engaged.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/childhoods_end/">
      <img src="https://alexgude.com/books/covers/childhoods_end.jpg" alt="Book cover of Childhood&#39;s End." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/childhoods_end/">
      <strong><cite class="book-title">Childhood’s End</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/arthur_c_clarke/"><span class="author-name">Arthur C. Clarke</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Childhood’s End</cite> is a classic sci-fi novel by <span class="author-name">Arthur C. Clarke</span>. It is about first contact
between humans and the mysterious Overlords, and the end of the human race.</p>
    </div>
  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="book-reviews" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[In 2024, I read 46 books and joined a sci-fi book club. From sequels to classics, here are my favorite reads of the year.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/book-reviews/vertical_books_from_the_marburg_image_archive.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/book-reviews/vertical_books_from_the_marburg_image_archive.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Steelcase Gesture Review: A Disappointing Upgrade</title>
      <link href="https://alexgude.com/blog/steelcase-gesture-review/" rel="alternate" type="text/html" title="Steelcase Gesture Review: A Disappointing Upgrade" />
      <published>2024-09-02T00:00:00-07:00</published>
      <updated>2024-09-02T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/steelcase_gesture_review</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/steelcase-gesture-review/"><![CDATA[<p>Although I have a <a href="/blog/nookdesk-review/">standing desk</a>, I still spend more than half
of my day sitting. For four years, I used the secondhand office chair my
brother-in-law gave me during the pandemic, but it was falling apart even
before I got it. So I decided it was time to upgrade.</p>

<p>I tried multiple different chairs in person, and although my heart wanted an
Aeron, my sit-test (for 10 minutes) told me I should get a <a href="https://store.steelcase.com/gesture">Steelcase
Gesture</a>. Giving me confidence was the fact that The Wirecutter
had rated it their <a href="https://www.nytimes.com/wirecutter/reviews/best-office-chair/">best chair</a>, saying “This is one of the most
adjustable chairs available—anyone can make it comfortable, regardless of
their height or size.” After eight months, I’ve decided they were wrong.</p>

<p>Overall, I am disappointed with the chair, which has an uncomfortable back,
poor recline, and arms that just don’t get out of the way enough to get close
to my desk. I have replaced it with an Aeron Remastered Size C.</p>

<h2 id="my-chair">My Chair</h2>

<p>I bought a Steelcase Gesture with:</p>

<ul>
  <li>
    <p>Headrest</p>
  </li>
  <li>
    <p>Upholstered Wrap Back</p>
  </li>
  <li>
    <p>Cogent Lizard Upholstery (5S94)</p>
  </li>
  <li>
    <p>Platinum Metallic Frame</p>
  </li>
  <li>
    <p>Matching Base</p>
  </li>
  <li>
    <p><span class="nowrap unit">360<abbr class="unit-abbr" title="Degrees">°</abbr></span> Arms</p>
  </li>
  <li>
    <p>Lumbar Support</p>
  </li>
  <li>
    <p>Wheels for Carpet</p>
  </li>
</ul>

<p>The retail price for this chair is $1679, but I got it through my work’s
office supplier for $1079.</p>

<h2 id="headrest">Headrest</h2>

<p>I debated whether or not I wanted a headrest and finally decided to get one.
You can mostly fold it out of the way if you don’t like it, so I figured I’d
rather give it a shot. It’s fine. I use it when lounging but not when sitting
normally. I would not get it again because it is superfluous.</p>

<h2 id="arms">Arms</h2>

<p>One of the selling points of the Gesture is how flexible the arms are. They
are very easy to move, easier than the Aeron’s, but I didn’t find this to be a
great selling point. Generally, I set the arms once and leave them forever. I
don’t adjust them for each new position or task as Steelcase assumes.</p>

<p>Additionally, the arms aren’t as nice as the Aeron’s. The Gesture’s arms don’t
go as low, about 6 1/2 inches off the seat compared to 5 1/2 for the Aeron.
This is extremely important to me because I have very long arms compared to my
torso; the Gesture’s arms force my shoulders up. The Gesture’s arms also stick
out further when pushed all the way back (10 inches from the back of the
chair, compared to 8 for the Aeron). This makes it harder to pull the chair
under a desk to get close to the keyboard. The pads are a little thinner on
the Gesture, making it slightly less comfortable.</p>

<h2 id="the-back">The Back</h2>

<p>The Gesture has a very upright back. It forces me into a very straight
position, even though I normally prefer a little bit of recline. The back
moves in two ways: it tilts top to bottom to try to adjust to the curve of
your spine, and it reclines.</p>

<p>The back is also very tall, extending about 24 inches above the seat (not
including the headrest). Because it is very straight, it makes contact with my
shoulder blades in a way I found uncomfortable. For the first week of use, I
got upper back pain which I assumed was just my body adjusting to “sitting
correctly”. Eight months later and I’m not so sure; the upper-back pain
persists (although far less severe) and my doctor thinks my posture is most of
the problem. I had no problem before using the Gesture.</p>

<h3 id="recline">Recline</h3>

<p>The Gesture’s back can be locked into 4 different positions, and the spring
tension can be adjusted from “you can’t push this back if you tried” to
“absolutely no resistance”. Unfortunately, the recline just doesn’t feel very
good on my chair. It is either too stiff or too loose. There is no middle
ground that feels supportive. I think my spring is broken because the Gestures
we have at my company’s office have better tension control and a more
comfortable recline.</p>

<p>The back moves separate from the seat, which feels OK. On the Aeron, the seat
slides a bit when you recline which feels more natural. Overall, the recline
on the Aeron feels much better.</p>

<h3 id="lumbar">Lumbar</h3>

<p>The lumbar support is <strong>TERRIBLE</strong>. It pokes me in the back and feels like
it’s trying to push me out of the chair. The curve of the backrest is already
aggressive, and the lumbar support makes it more so. There is an <a href="https://www.reddit.com/r/OfficeChairs/comments/rixars/psa_remove_lumbar_support_from_your_steelcase/">entire
thread on Reddit</a><sup style="anchor-name:--fnref-reddit" id="fnref:reddit"><a href="#fn:reddit" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> about how the lumbar support ruins the
chair. I followed <a href="/files/steelcase//steelcase_gesture_back_replacement_guide-939544221revA.pdf">Steelcase’s guide to removing the backrest pad</a> to
take out the lumbar support and it helped a lot (but I still didn’t find it
comfortable). I would not buy the lumbar support again.</p>

<h2 id="the-seat">The Seat</h2>

<p>The seat depth is adjustable. Because I’m tall (6 foot 1 inch), I have the
seat all the way forward. This leaves a small gap between the back of the seat
and the backrest. The height is also adjustable, and I can make it high enough
to fit me.</p>

<p>Some people complain that the woven and padded seat retains heat and isn’t
comfortable on hot days. I used the chair during the summer and didn’t notice
heat problems, even when it was <span class="nowrap unit">110 <abbr class="unit-abbr" title="Degrees Fahrenheit">°F</abbr></span> outside and a
little over <span class="nowrap unit">80 <abbr class="unit-abbr" title="Degrees Fahrenheit">°F</abbr></span> in my office.</p>

<h2 id="final-thoughts">Final Thoughts</h2>

<p>I don’t like my Steelcase Gesture. For eight months, walking into my home
office each morning came with a sinking feeling, knowing I’d have to sit in
this chair. I replaced it with an Aeron, and the difference is night and
day—I couldn’t be happier with the switch.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:reddit">

      <p><figure class="cited-quote"><blockquote cite="https://www.reddit.com/r/OfficeChairs/comments/rixars/psa_remove_lumbar_support_from_your_steelcase/"> <p>So I got a new Steelcase Gesture and for about a week couldn’t understand why does this cost $1k+ and how someone can possibly prefer it to Aeron. I even started to get the back pain in the middle of the spine - something that has never happened to me in my seating life ever.</p>  <p>Then I read here about Lumbar support being in the way of their flexible back. I could not really find the right position for the lubmar so I figured ok I will just remove it, if it doesn’t help I will sell the damn thing at loss asap.</p>  <p>So without the Lumbar Support Gesture is <strong>amazing</strong>. The back flexes instead of cutting into your spine, you can lean on the chair and if you touch it behind the backrest you would notice that the design has the ribs that are supposed to have some give.</p>  <p>The lumbar support itself is fairly rigid piece of plastic that, frankly, feels alien inside the chair. Once I took it out, I felt a bit like I did some life-savingv surgery on the poor thing.</p>  <p>I have no idea who would prefer to have the lumbar support in it, especially compared to how good it is without it. I’d say with lumbar support it’s a 6.5 chair (and, considering the price for the new one, more like 5.5). Without it’s a solid 8.5-9. My only gripe now is that it’s on the hot side, I wish it had a mesh seat, but I think I can survive, the adjustable arms and the overall smoothness of the experience without the lumbar plate is worth it.</p>  <p>I had a headrest version, and the removal procedure isn’t exactly trivial and takes some force, but took me about 30 minutes. If you have the wrapped back be prepared to take some risks, inserting the lever underneath it to pop it off.</p>  <p>After that you screw off 4 screws (torx) and slide up the front seating pad. Then you’d need to carefully slide out the lumbar plate (I did not have to disconnect anything there, just the textile in a couple of places).</p>  <p>That’s it.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">myreptilianbrain. <a href="https://www.reddit.com/r/OfficeChairs/comments/rixars/psa_remove_lumbar_support_from_your_steelcase/">“PSA: Remove lumbar support from your Steelcase Gesture”</a> <cite>Reddit, r/OfficeChairs</cite>. 2021-12-18.</span></figcaption></figure> <a href="#fnref:reddit" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[I spent eight months with the highly-rated Steelcase Gesture, only to be disappointed. In this review, I break down why this expensive ergonomic chair didn't work for me.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/steelcase/steelcase_gesture_with_headrest_360_arms_in_lizard.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/steelcase/steelcase_gesture_with_headrest_360_arms_in_lizard.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Favorite Books of 2023</title>
      <link href="https://alexgude.com/blog/favorite-books-of-2023/" rel="alternate" type="text/html" title="My Favorite Books of 2023" />
      <published>2024-01-01T00:00:00-08:00</published>
      <updated>2024-01-01T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/favorite_books_of_2023</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/favorite-books-of-2023/"><![CDATA[<p>Social media died for me in 2023. Twitter and Reddit both shut down their third-party
APIs, making accessing them a pain, and the quality of Twitter decreased
significantly with its sale to Elon Musk. I realized how much time I had
wasted on social media and decided to spend it more wisely by reading books. I
bought a Kindle in October and finished 20 books between then and the end of
the year.</p>

<p>I <a href="/books/">review all the books I read</a>, here is a selection of my
favorites that I read in 2023:</p>

<h2 id="blindsight-by-peter-watts"><cite class="book-title">Blindsight</cite> by <span class="author-name">Peter Watts</span></h2>

<p><cite class="book-title">Blindsight</cite> does a great job of exploring the
nature of consciousness and intelligence. Watts keeps the tension high and the
plot moving quickly in this thought-provoking sci-fi novel. My favorite book
of the year!</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/blindsight/">
      <img src="https://alexgude.com/books/covers/blindsight.jpg" alt="Book cover of Blindsight." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/blindsight/">
      <strong><cite class="book-title">Blindsight</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_watts/"><span class="author-name">Peter Watts</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Blindsight</cite> is a hard sci-fi novel about first contact with
aliens in the near future. A crew of four transhumans and a vampire are sent
on a spaceship to investigate an anomaly in the solar system after a swarm of
alien probes scan Earth.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-chronicle-of-the-fallers-by-peter-f-hamilton"><span class="book-series">The Chronicle of the Fallers</span> by <span class="author-name">Peter F. Hamilton</span></h2>

<p>Hamilton is known for his space opera, but <cite class="book-title">The Abyss
Beyond Dreams</cite> is more urban fantasy set during the Russian Revolution
(in space) and <cite class="book-title">Night Without Stars</cite> is a
thriller set during the Cold War (again, in space). Both feature Commonwealth
citizens with special knowledge as <em>“Outside Context Problems”</em>, pulling the
stories into science fiction territory.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_abyss_beyond_dreams/">
      <img src="https://alexgude.com/books/covers/the_abyss_beyond_dreams.jpg" alt="Book cover of The Abyss Beyond Dreams." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_abyss_beyond_dreams/">
      <strong><cite class="book-title">The Abyss Beyond Dreams</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_f_hamilton/"><span class="author-name">Peter F. Hamilton</span></a></span>
    <div class="book-rating star-rating-3" role="img" aria-label="Rating: 3 out of 5 stars. Good — I liked it" title="Good — I liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Abyss Beyond Dreams</cite> starts off <span class="book-series">The Chronicle of the Fallers</span>, another series in <span class="author-name">Peter F. Hamilton</span>’s Commonwealth universe. Though billed as space opera, it often reads more as
urban fantasy since most of the story occurs on the planet Bienvenido inside
the Void where steam engines are their most advanced technology.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/night_without_stars/">
      <img src="https://alexgude.com/books/covers/night_without_stars.jpg" alt="Book cover of Night Without Stars." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/night_without_stars/">
      <strong><cite class="book-title">Night Without Stars</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_f_hamilton/"><span class="author-name">Peter F. Hamilton</span></a></span>
    <div class="book-rating star-rating-3" role="img" aria-label="Rating: 3 out of 5 stars. Good — I liked it" title="Good — I liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Night Without Stars</cite> is the second book in <span class="book-series">The Chronicle of the Fallers</span>. It is action packed, with great pacing, and complex characters.
It is my new favorite <span class="author-name">Peter F. Hamilton</span> book.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-fall-of-hyperion-by-dan-simmons"><cite class="book-title">The Fall of Hyperion</cite> by <span class="author-name">Dan Simmons</span></h2>

<p>I enjoyed the sequel to <a href="/books/hyperion/"><cite class="book-title">Hyperion</cite></a> the most of the two books
because it tied the personal story of the pilgrims to a much broader galactic
conflict. Interestingly, you can see a lot of ideas in the Hyperion Cantos
that Hamilton later adopted in his Commonwealth Saga including wormholes, a
breakaway-but-helpful AI, and different factions of scheming AI who either
want to eradicate the humans or uplift them.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/the_fall_of_hyperion/review-2023-10-27/">
      <img src="https://alexgude.com/books/covers/the_fall_of_hyperion.jpg" alt="Book cover of The Fall of Hyperion." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/the_fall_of_hyperion/review-2023-10-27/">
      <strong><cite class="book-title">The Fall of Hyperion</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/dan_simmons/"><span class="author-name">Dan Simmons</span></a></span>
    <div class="book-rating star-rating-5" role="img" aria-label="Rating: 5 out of 5 stars. Masterpiece — I loved it" title="Masterpiece — I loved it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">The Fall of Hyperion</cite> is a sequel that outshines its predecessor. It is
everything I was expecting from <a href="/books/hyperion/"><cite class="book-title">Hyperion</cite></a> and more! A true
masterpiece.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="the-commonwealth-saga-by-peter-f-hamilton"><span class="book-series">The Commonwealth Saga</span> by <span class="author-name">Peter F. Hamilton</span></h2>

<p>Epic space opera with a massive cast of characters and incredible pacing.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/pandoras_star/">
      <img src="https://alexgude.com/books/covers/pandoras_star.jpg" alt="Book cover of Pandora&#39;s Star." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/pandoras_star/">
      <strong><cite class="book-title">Pandora’s Star</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_f_hamilton/"><span class="author-name">Peter F. Hamilton</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p>I couldn’t put <cite class="book-title">Pandora’s Star</cite> down! It is a sci-fi book that reads
more like a thriller. There were always new mysteries that just a few more
pages promised the answers to.</p>
    </div>
  </div>
</li>
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/judas_unchained/">
      <img src="https://alexgude.com/books/covers/judas_unchained.jpg" alt="Book cover of Judas Unchained." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/judas_unchained/">
      <strong><cite class="book-title">Judas Unchained</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/peter_f_hamilton/"><span class="author-name">Peter F. Hamilton</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p>The sequel to <a href="/books/pandoras_star/"><cite class="book-title">Pandora’s Star</cite></a>, <cite class="book-title">Judas Unchained</cite> continues right where the last one left off, but with the
action ramped up to 11. The various storylines and loose threads come together
one by one until it’s the good guys racing against the bad guys for the fate
of the universe.</p>
    </div>
  </div>
</li>
</ul>

<h2 id="serpent-valley-by-scott-warren"><cite class="book-title">Serpent Valley</cite> by <span class="author-name">Scott Warren</span></h2>

<p>1980s mech sci-fi re-imagined for the 21st century. Warren’s self-published
series takes a few books to really find its feet, but once it does, it’s a
quick, fun, nostalgic read. The third book, <cite class="book-title">Serpent
Valley</cite>, exemplifies the series.</p>

<ul class="card-grid">
<li class="book-card">
  <div class="card-element card-book-cover">
    <a href="https://alexgude.com/books/serpent_valley/">
      <img src="https://alexgude.com/books/covers/serpent_valley.jpg" alt="Book cover of Serpent Valley." />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/books/serpent_valley/">
      <strong><cite class="book-title">Serpent Valley</cite></strong>
    </a>
    <span class="by-author"> by <a href="/books/authors/scott_warren/"><span class="author-name">Scott Warren</span></a></span>
    <div class="book-rating star-rating-4" role="img" aria-label="Rating: 4 out of 5 stars. Great — I really liked it" title="Great — I really liked it"><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star full_star" aria-hidden="true">★</span><span class="book_star empty_star" aria-hidden="true">☆</span></div>
    <div class="card-element card-text">
      <p><cite class="book-title">Serpent Valley</cite>, the third book in the <span class="book-series">War Horses</span> series, is another quick, action-packed read—but without the flaws
holding back its predecessors. Easily my favorite of the series so far!</p>
    </div>
  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="book-reviews" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[After abandoning social media in 2023, I read books instead. Read on for my favorites.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/book-reviews/corner_books_from_the_marburg_image_archive.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/book-reviews/corner_books_from_the_marburg_image_archive.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Using Large Language Models To Clean Data</title>
      <link href="https://alexgude.com/blog/large-language-models-for-data-cleaning/" rel="alternate" type="text/html" title="Using Large Language Models To Clean Data" />
      <published>2023-10-15T00:00:00-07:00</published>
      <updated>2023-10-15T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/large_language_models_for_data_cleaning</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/large-language-models-for-data-cleaning/"><![CDATA[<p>I maintain the <a href="/blog/switrs-to-sqlite/">SWITRS-to-sqlite</a> Python library that parses and cleans
up California Highway Patrol’s traffic collision database. One of the fields
the responding officer has to fill out at the scene of the crash is the
make<sup style="anchor-name:--fnref-make" id="fnref:make"><a href="#fn:make" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> of the vehicle. This field is a free text field, but there is a
relatively small number of common brands, so it should be mapped to a
categorical column.</p>

<p>This is straightforward when the officer writes <code class="language-plaintext highlighter-rouge">FORD</code> or <code class="language-plaintext highlighter-rouge">HONDA</code>, which they
mostly do. But since the officer can write anything, they occasionally make it
a little harder on us by abbreviating or mistyping, for example <code class="language-plaintext highlighter-rouge">VOLX</code> and
<code class="language-plaintext highlighter-rouge">DODDGE</code>. And sometimes they make it impossible by writing <code class="language-plaintext highlighter-rouge">--</code> or <code class="language-plaintext highlighter-rouge">______</code>.</p>

<p>The solution is to go through, one by one, and create a <a href="/blog/python-patterns-enum/">mapping</a>
like:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Enumeration of common vehicle makes
</span><span class="nd">@unique</span>
<span class="k">class</span> <span class="nc">Make</span><span class="p">(</span><span class="n">Enum</span><span class="p">):</span>
  <span class="n">CHEVROLET</span>  <span class="o">=</span> <span class="sh">"</span><span class="s">chevrolet</span><span class="sh">"</span>
  <span class="n">GMC</span>        <span class="o">=</span> <span class="sh">"</span><span class="s">gmc</span><span class="sh">"</span>
  <span class="n">HINO</span>       <span class="o">=</span> <span class="sh">"</span><span class="s">hino</span><span class="sh">"</span>
  <span class="n">INFINITI</span>   <span class="o">=</span> <span class="sh">"</span><span class="s">infiniti</span><span class="sh">"</span>
  <span class="n">MITSUBISHI</span> <span class="o">=</span> <span class="sh">"</span><span class="s">mitsubishi</span><span class="sh">"</span>
  <span class="c1"># Special Token for unknown make
</span>  <span class="n">NONE</span>       <span class="o">=</span> <span class="bp">None</span>

<span class="c1"># Dictionary mapping raw values to Make enum
</span><span class="n">make_map</span> <span class="o">=</span> <span class="p">{</span>
  <span class="sh">"</span><span class="s">CHEVRLT</span><span class="sh">"</span><span class="p">:</span>  <span class="n">Make</span><span class="p">.</span><span class="n">CHEVROLET</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">HINO/</span><span class="sh">"</span><span class="p">:</span>    <span class="n">Make</span><span class="p">.</span><span class="n">HINO</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">INFINITY</span><span class="sh">"</span><span class="p">:</span> <span class="n">Make</span><span class="p">.</span><span class="n">INFINITI</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">MITSUB</span><span class="sh">"</span><span class="p">:</span>   <span class="n">Make</span><span class="p">.</span><span class="n">MITSUBISHI</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">TAHOE</span><span class="sh">"</span><span class="p">:</span>    <span class="n">Make</span><span class="p">.</span><span class="n">GMC</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">UKNOWN</span><span class="sh">"</span><span class="p">:</span>   <span class="n">Make</span><span class="p">.</span><span class="n">NONE</span><span class="p">,</span>
<span class="p">}</span>
</code></pre></div></div>

<p>As someone who did this mapping by hand for <a href="https://github.com/agude/SWITRS-to-SQLite/blob/85ac7e7850680bd47f3fef5a44ab180d8ee9dd8b/switrs_to_sqlite/make_map.py">over 900 entries</a>, it is
quite tedious. Fortunately, making sense of mangled text is something <a href="/blog/good-uses-for-large-language-models/">Large
Language Models (LLMs) are pretty good at</a>!</p>

<h2 id="automating">Automating</h2>

<p>The goal is to perform few-shot, multi-label classification of vehicle makes.
Few-shot because we are going to give the model just a handful of examples of
what output we expect, and multi-label because there are many possible vehicle
makes it will have to map to.</p>

<h3 id="prompting">Prompting</h3>

<p>The first step is to write a prompt explaining the task to the model, the
expected return value, and a few examples of input and correct outputs. Here
is a shortened version, the full one is <a href="/blog/llm-data/prompt/">here</a>, starting with the
instructions:</p>

<div class="chatgpt-edit-block"> 
<div class="chatgpt-prompt-only">
    <blockquote>
      <p>I am working with a dataset of traffic collisions from California. One of
the fields is the “make” of the vehicle, for example, “Honda”, “Ford”,
“Peterbilt”, etc.</p>

      <p>But this field a free-text field filled out by the CHP officer on the scene
of the collision. As such there are misspellings, abbreviations, and other
mistakes that have to be fixed.</p>

      <p>I have created a set of makes as follows (including <code class="language-plaintext highlighter-rouge">NONE</code> as a placeholder
for unknown values). Here is the list in a Python <code class="language-plaintext highlighter-rouge">Enum</code>:</p>

      <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@unique</span>
<span class="k">class</span> <span class="nc">Make</span><span class="p">(</span><span class="n">Enum</span><span class="p">):</span>
    <span class="n">ACADIAN</span>                 <span class="o">=</span> <span class="sh">"</span><span class="s">acadian</span><span class="sh">"</span>
    <span class="n">ACURA</span>                   <span class="o">=</span> <span class="sh">"</span><span class="s">acura</span><span class="sh">"</span>
    <span class="n">ALFA_ROMERO</span>             <span class="o">=</span> <span class="sh">"</span><span class="s">alfa romera</span><span class="sh">"</span>
    <span class="n">AMC</span>                     <span class="o">=</span> <span class="sh">"</span><span class="s">american motors</span><span class="sh">"</span>
    <span class="bp">...</span>
</code></pre></div>      </div>

      <p>Take note that anything unknown should be tagged with <code class="language-plaintext highlighter-rouge">Make.None</code>. And do
not make up new Enum values.</p>
    </blockquote>
  </div>
</div>

<p>Then the output format, with instructions to include an explanation of its
logic first, which can <a href="https://arxiv.org/abs/2201.11903">help model accuracy</a>:</p>

<div class="chatgpt-edit-block"> 
<div class="chatgpt-prompt-only">
    <blockquote>
      <p>I will provide you with a string. You are to return a Python dictionary with
the following keys, in this same order:</p>

      <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span>
  <span class="n">explanation</span><span class="p">:</span> <span class="sh">"</span><span class="s">An explanation of why you think the enum value is a good match, or why there is no match possible.</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">input_string</span><span class="p">:</span> <span class="sh">"</span><span class="s">The input string</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">enum</span><span class="p">:</span> <span class="sh">"</span><span class="s">The correct enum from above</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">no_match</span><span class="p">:</span> <span class="sh">"</span><span class="s">`True` or `False`. True if there is no matching enum or no way to make a match, otherwise False.</span><span class="sh">"</span><span class="p">,</span> 
<span class="p">}</span>
</code></pre></div>      </div>
    </blockquote>
  </div>
</div>

<p>And finally some examples of inputs and correct outputs:</p>

<div class="chatgpt-edit-block"> 
<div class="chatgpt-prompt-only">
    <blockquote>
      <p>For example, for the input <code class="language-plaintext highlighter-rouge">VOLX</code>:</p>

      <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span>
  <span class="n">explanation</span><span class="p">:</span> <span class="sh">"""</span><span class="s">VOLX is pronouced similarly to </span><span class="sh">'</span><span class="s">Volks</span><span class="sh">'</span><span class="s"> and therefore this is
    probably an abbreviation of </span><span class="sh">'</span><span class="s">Volkswagen</span><span class="sh">'</span><span class="s">. There is an enum value for
    Volkswagon, `Make.VOLKSWAGEN`, already so we use that.</span><span class="sh">"""</span><span class="p">,</span>
  <span class="n">input_string</span><span class="p">:</span> <span class="sh">"</span><span class="s">VOLX</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">enum</span><span class="p">:</span> <span class="n">make</span><span class="p">.</span><span class="n">VOLKSWAGEN</span><span class="p">,</span>
  <span class="n">no_match</span><span class="p">:</span> <span class="bp">False</span><span class="p">,</span>
<span class="p">}</span>
</code></pre></div>      </div>
    </blockquote>
  </div>
</div>

<h3 id="answers">Answers</h3>

<p>Since I was manually copying the prompt into the model’s web interface, I used
batches of 100–200 string sorted alphabetically. With API access, I could
have used <a href="https://en.wikipedia.org/w/index.php?title=Prompt_engineering&amp;oldid=1179231833#Retrieval-augmented_generation">retrieval-augmented generation</a> to create custom examples for
each string while sending them one at a time.</p>

<p>Splitting the data into batches helped the model figure out very short
entries. For example, the model failed when given <code class="language-plaintext highlighter-rouge">WNBG</code> (Winnebago) by
itself, but succeeded when I gave it the list:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>WINN
WINNE
WINNEBAG
WINNEBAGO
WINNI
WNBG
WNBGO
</code></pre></div></div>

<p>I believe seeing multiple short versions next to each other helped the model
infer the right mapping.</p>

<h3 id="performance">Performance</h3>

<p>I obtained the following performance on my 902 hand-mapped entries:</p>

<ul>
  <li>
    <p>The model correctly fixed 2 entries that I had gotten wrong.</p>
  </li>
  <li>
    <p>It matched 682 (75.6%) of my hand-labeled mappings.</p>
  </li>
  <li>
    <p>It missed 218 (24.1%) of the mappings, frequently using made-up enum values.</p>
  </li>
</ul>

<p>This is reasonably good performance, as finding wrong entries is pretty quick
(and many could be fixed with find and replace).</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:make">

      <p>The “make” of a vehicle is the brand of the manufacturer, like ‘Honda’,
‘Ford’, ‘Tesla’, etc. <a href="#fnref:make" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
          <category term="california-traffic-data" />
        
      

      

      
      
        <summary type="html"><![CDATA[Manually fixing messy data is tedious and slow. But thankfully, LLMs are pretty good at piecing together mangled text. Read on to find out how!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/llm-data/00045-1994538970-a_simple_color_pencil_drawing_a_robot,_inspecting_a_car,_holding_a_clipboard,_white_background.png" />
        <media:content medium="image" url="https://alexgude.com/files/llm-data/00045-1994538970-a_simple_color_pencil_drawing_a_robot,_inspecting_a_car,_holding_a_clipboard,_white_background.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">A 1850-mile Review of the RadWagon 3</title>
      <link href="https://alexgude.com/blog/radwagon-long-term-review/" rel="alternate" type="text/html" title="A 1850-mile Review of the RadWagon 3" />
      <published>2023-07-24T00:00:00-07:00</published>
      <updated>2023-07-24T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/radwagon_long_term_review</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/radwagon-long-term-review/"><![CDATA[<p>My wife and I bought a RadWagon 3, an inexpensive electric cargo bike made by
<a href="https://www.radpowerbikes.com">Rad Power Bikes</a>, right before the pandemic hit in early 2020. We wanted
a cargo bike that we could use to haul kids around, and we decided on a cheap
one since we weren’t sure it would fit our lifestyle.</p>

<p>We’ve now put 1850 miles on the bike, mostly taking the kids to local parks.
After all the time spent with it, would we recommend it to another family?</p>

<p>In short: we <strong>love the utility an electric cargo bike offers</strong>, and I think
we will always have one in the garage to supplement our minivan, <strong>but the
RadWagon has some significant drawbacks</strong> that make recommending it
difficult. Read on for the full review.</p>

<h2 id="the-bike">The Bike</h2>

<p>We ordered the Rad Wagon from the Rad’s website for $1500. With a “caboose”
enclosure and pads for the kids adding $250, plus tax, our total was
$1901—not cheap, but almost a third of comparable cargo bikes which are in
the $4000–6000 range.</p>

<h3 id="assembly-and-adjustment">Assembly and Adjustment</h3>

<p>The bike comes in a box and you must assemble it yourself. The assembly was
not too hard, but I also have a lot of experience with bikes and bike
maintenance which helped.</p>

<p>The seat post and headset are highly adjustable, allowing people of varying
heights to ride the RadWagon. Both my wife (5’4”) and I (6’1”) can ride it
comfortably, but we’re both near the limits. I have the seat all the way up
and she has it all the way down.</p>

<h3 id="motor">Motor</h3>

<p>The 750W rear hub motor easily brings the bike to its 20 miles per hour
computer-limited top speed. The bike has a throttle, which I love for getting
started from a stop, and enough power to carry me, two kids, and some balance
bikes up a steep hill. The downside though is the hub motor puts a lot of
stress on the rear wheel.</p>

<h3 id="flaws">Flaws</h3>

<p>A flaw in the design is the rear brake. The RadWagon uses cheap, mechanical
disc brakes, which are enough to stop the bike when they’re well aligned, but
which need constant attention to keep them that way. The motor blocks the
typical through-spoke access for adjusting the rear brake. Instead, Rad makes
a special flat Allen wrench that fits between the brake and motor but
adjusting remains hard.</p>

<p>A major flaw is the rear spokes. They are stressed by both the motor—which
puts all its power through the rear wheel—and the cargo. The spokes were not
tight enough from the factory and I broke several in the first 200 miles. I
have broken fewer since replacing and retightening, but I still break one
periodically, which is annoying for me and likely a dealbreaker for less
experienced riders.</p>

<h2 id="support">Support</h2>

<p>I’ve contacted Rad’s support several times—to order spokes and the brake
tool, and to replace a faulty accessory. They were generally quick and helpful
but I’ve never needed support for other bikes. And the last interaction was
terrible: Rad’s front basket was defective, and after I sent them photographic
proof, they accused me of being unable to use a screw driver and stopped
responding. It was the worst support experience I’ve ever had.</p>

<h2 id="final-thoughts">Final Thoughts</h2>

<p>The RadWagon was a savior during the pandemic, letting us escape the house and
ride when we’d otherwise be trapped inside. It also beats driving kids in a
car—hauling them on the back makes getting to the park part of the fun.</p>

<p>But the bike needs constant maintenance that is difficult even for an
experienced mechanic, and Rad’s support is not great. I wish I had purchased a
higher-quality bike that wouldn’t fail so frequently.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[The RadWagon electric cargo bike was a savior during lockdown family rides, but frustrating maintenance and support issues disappoint. Read on for my in-depth review.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/radwagon/loaded_radwagon_3.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/radwagon/loaded_radwagon_3.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Machine Learning Deployment: Return Actions, Not Scores</title>
      <link href="https://alexgude.com/blog/machine-learning-return-actions-not-scores/" rel="alternate" type="text/html" title="Machine Learning Deployment: Return Actions, Not Scores" />
      <published>2023-06-26T00:00:00-07:00</published>
      <updated>2023-06-26T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/machine_learning_return_actions_not_scores</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/machine-learning-return-actions-not-scores/"><![CDATA[<p>At a previous job, my team built models that stopped ATOs—<a href="https://en.wikipedia.org/wiki/Credit_card_fraud#Account_takeover">Account
takeovers</a>, where a fraudster steals someone’s account credentials
and attempts to use them. The engineering team that owned the login flow would
call our model <a href="https://en.wikipedia.org/wiki/API">API</a>, and we would return the model score. The
engineering team had a threshold in their code, and if the score crossed that
threshold, they would take some action.</p>

<p>You can probably already see the problem: APIs are <strong>meant to hide the inner
workings behind them</strong>. But by returning the raw model scores, we revealed too
much detail. Any changes to the model, like retraining it, could change the
scores and break the front end.</p>

<p>In my guide to <a href="/blog/machine-learning-deployment-shadow-mode/#in-front-of-the-api"><em>deploying machine learning models in shadow
mode</em></a>, I stated that deploying changes “in front of the
API” has the advantage of giving the calling team control. This is precisely
why we built the ATO API the way we did: to address the organizational issue
that the engineering team did not trust the machine learning team.</p>

<p>But if your teams trust each other, there is a much better way to build.</p>

<h2 id="what-is-a-better-way">What is a better way?</h2>

<p>A better way is for the API to return <strong>a set of actions</strong>. For example, the
ATO model API might return the following actions:</p>

<ul>
  <li>
    <p><em>Allow</em>: The login looks fine, allow it.</p>
  </li>
  <li>
    <p><em>Step-up</em>: The login looks odd, require the user to provide a second factor
of authentication, such as a code sent to their email.</p>
  </li>
  <li>
    <p><em>Lock</em>: The login looks clearly fraudulent, deny the login and lock the
account until the user recovers it.</p>
  </li>
</ul>

<p>These actions do a really good job of hiding the implementation behind the
API. You can freely change thresholds when the model performance changes,
retrain the model, or even replace it entirely.</p>

<p>But you can do something else too, you can add more models!</p>

<h3 id="using-multiple-systems">Using multiple systems</h3>

<p>A common fraud-prevention strategy is to train a model for each new fraud
pattern identified. This allows each model to be highly <a href="https://en.wikipedia.org/wiki/Precision_and_recall">precise</a>,
while also improving the <a href="https://en.wikipedia.org/wiki/Precision_and_recall">recall</a> of the overall system. These
multi-model systems are often augmented with simple rules, such as “No logins
from Russia allowed.” In the end, the system takes the outputs of the various
models and rules and aggregates them in some way. In our ATO example, the
system returns the most drastic action recommended by any model or rule.</p>

<p>In code:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">ato_api</span><span class="p">(</span><span class="n">event_token</span><span class="p">):</span>
  <span class="c1"># List of actions returned by all the models and rules,
</span>  <span class="c1"># consists of values from {'Allow', 'Step-up', 'Lock'}
</span>  <span class="n">all_results</span> <span class="o">=</span> <span class="nf">get_ato_system_results</span><span class="p">(</span><span class="n">event_token</span><span class="p">)</span>

  <span class="k">if</span> <span class="sh">'</span><span class="s">Lock</span><span class="sh">'</span> <span class="ow">in</span> <span class="n">all_results</span><span class="p">:</span>
    <span class="k">return</span> <span class="sh">'</span><span class="s">Lock</span><span class="sh">'</span>
  <span class="k">elif</span> <span class="sh">'</span><span class="s">Step-up</span><span class="sh">'</span> <span class="ow">in</span> <span class="n">all_results</span><span class="p">:</span>
    <span class="k">return</span> <span class="sh">'</span><span class="s">Step-up</span><span class="sh">'</span>

  <span class="k">return</span> <span class="sh">'</span><span class="s">Allow</span><span class="sh">'</span>
</code></pre></div></div>

<p>Of course, this is a great place to use <a href="/blog/python-patterns-enum/">enums</a> and
<a href="/blog/python-patterns-max-not-if/">max</a>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">enum</span> <span class="kn">import</span> <span class="n">IntEnum</span><span class="p">,</span> <span class="n">unique</span>

<span class="nd">@unique</span>
<span class="k">class</span> <span class="nc">Action</span><span class="p">(</span><span class="n">IntEnum</span><span class="p">):</span>
  <span class="n">ALLOW</span> <span class="o">=</span> <span class="mi">0</span>
  <span class="n">STEPUP</span> <span class="o">=</span> <span class="mi">1</span>
  <span class="n">LOCK</span> <span class="o">=</span> <span class="mi">2</span>

<span class="k">def</span> <span class="nf">ato_api</span><span class="p">(</span><span class="n">event_token</span><span class="p">):</span>
  <span class="c1"># List of actions returned by all the models and rules,
</span>  <span class="c1"># consists of values from Action() enum
</span>  <span class="n">all_results</span> <span class="o">=</span> <span class="nf">get_ato_system_results</span><span class="p">(</span><span class="n">event_token</span><span class="p">)</span>

  <span class="k">return</span> <span class="nf">max</span><span class="p">(</span><span class="n">all_results</span><span class="p">)</span>
</code></pre></div></div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="machine-learning" />
        
          <category term="machine-learning-engineering" />
        
      

      

      
      
        <summary type="html"><![CDATA[A poorly designed machine learning model API will leave you trapped. Properly hiding your implementation will make life much easier!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/ml-api/00678-3489016266-a_simple_color_pencil_drawing_a_cute_robot,_plugging_cat5_cable_into_a_network_switch,_white_background.png" />
        <media:content medium="image" url="https://alexgude.com/files/ml-api/00678-3489016266-a_simple_color_pencil_drawing_a_cute_robot,_plugging_cat5_cable_into_a_network_switch,_white_background.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Claude Solves SAT Analogies</title>
      <link href="https://alexgude.com/blog/claude-sat-analogies/" rel="alternate" type="text/html" title="Claude Solves SAT Analogies" />
      <published>2023-05-29T00:00:00-07:00</published>
      <updated>2023-05-29T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/claude_sat_analogies</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/claude-sat-analogies/"><![CDATA[<p>Several years ago, I <a href="/blog/sat2vec/">tried to get Word2Vec to solve SAT
analogies</a>. It did not go well. Word2Vec got just 8 out of 36
right.</p>

<p>But in the last 7 years language models have gotten much, <strong>MUCH</strong> better. I
wondered how a state-of-the-art model, one too large to run on my computer,
would perform on the same questions.</p>

<p>To find out, I ran the analogies through <a href="https://www.anthropic.com/">Anthropic’s</a> biggest
model: <a href="https://www.anthropic.com/index/introducing-claude">Claude</a>.</p>

<h2 id="experimental-setup">Experimental Setup</h2>

<p>I gave Claude the following instructions:</p>

<div class="chatgpt-edit-block">
<div class="chatgpt-prompt-only">
    <blockquote>
      <p>We’re going to solve SAT analogy questions. I’ll give you a pair of words
like:</p>

      <p>“authenticity : counterfeit”</p>

      <p>And you determine the relationship between the two words, and then pick the
pair from the next 5 with the same relation. So in this case I would give
you:</p>

      <p>reliability : erratic</p>

      <p>mobility : energetic</p>

      <p>argument : contradictory</p>

      <p>reserve : reticent</p>

      <p>anticipation : solemn</p>

      <p>You need to pick:</p>

      <p>“reliability : erratic”</p>

      <p>Make sure to explain the relationship and why you chose the answer you did.</p>
    </blockquote>
  </div>
</div>

<p>Then I gave it an analogy like:</p>

<div class="chatgpt-edit-block">
<div class="chatgpt-prompt-only">
    <blockquote>
      <p>fix : stable</p>

      <p><span>------------------------</span></p>

      <p>fortify : strong</p>

      <p>captivate : attractive</p>

      <p>deter : active</p>

      <p>furrow : productive</p>

      <p>mend : torn</p>
    </blockquote>
  </div>
</div>

<details>
  <summary>
    <p>And it would reply with its answer and reasoning (click “Show” to see all of
it):</p>

    <div class="chatgpt-edit-block">
<div class="chatgpt-output-only">
        <blockquote>
          <p>Okay, for this analogy:</p>

          <p>agenda : conference</p>

          <p>The relationship here is:</p>

          <p>An agenda is a list that provides the plan or schedule for a conference. So
a conference relies on and uses an agenda.</p>

          <p>Based on this, the choice that shows a similar relationship is:</p>

          <p>map : trip</p>
        </blockquote>
      </div>
</div>

  </summary>

  <div class="chatgpt-edit-block">
<div class="chatgpt-output-only">
      <blockquote>
        <p>Okay, for this analogy:</p>

        <p>agenda : conference</p>

        <p>The relationship here is:</p>

        <p>An agenda is a list that provides the plan or schedule for a conference. So
a conference relies on and uses an agenda.</p>

        <p>Based on this, the choice that shows a similar relationship is:</p>

        <p>map : trip</p>

        <p>Because a trip relies on and uses a map. A map provides the plan and
guidance for a trip, just as an agenda does for a conference.</p>

        <p>The other choices do not show this same kind of relationship:</p>

        <ul>
          <li>Teacher : class - A teacher facilitates a class but a class does not rely
on or use a teacher in the same direct way that a conference uses an
agenda or a trip uses a map.</li>
          <li>Executive : employee - An executive oversees employees but employees do
not directly rely on or make use of an executive the same way a conference
uses an agenda or a trip uses a map.</li>
          <li>Man : woman - These do not have a direct facilitator/user relationship.
They do not rely on each other in the same way.</li>
          <li>Agency : assignment - While an agency may give out assignments, an
assignment does not directly rely on or use an agency in the way a
conference uses an agenda or a trip uses a map.</li>
        </ul>

        <p>So by analyzing each choice, map : trip is the one option that shows the
closest relationship to the original pair’s dynamic of agenda : conference.
Please let me know if this makes sense or if you have another analogy
challenge for me!</p>
      </blockquote>
    </div>
</div>
</details>

<p>One potential issue: I left the analogies and their possible answers in the
sorted order from my word2vec post. This means the analogies word2vec got
right were up first, with the correct answer in the first position. Then the
ones word2vec had the correct answer second, and so on. I randomized a few of
the analogies to check if this was a significant problem, and it did not seem
to be. If I had API access, I would have randomized the order of all the
potential answers and run multiple tests.</p>

<h2 id="results">Results</h2>

<p>Unsurprisingly, Claude did much better than Word2Vec. It got 25 out of 35
analogies correct<sup style="anchor-name:--fnref-previous" id="fnref:previous"><a href="#fn:previous" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> compared to 8 out of 36 for Word2Vec.</p>

<p>The most frequent failure appears to be Claude identifying the right
relationship, but then incorrectly applying it to choose the answer. For
example this analogy (correct answer <strong>bold</strong>, Claude’s answer in <u style="color:ForestGreen">green
underline</u>):</p>

<div class="low-width-table" style="max-width: 30%">

  <table>
    <thead>
      <tr>
        <th style="text-align: left">hospitable : courtesy</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td style="text-align: left"><strong>vindictive : spite</strong></td>
      </tr>
      <tr>
        <td style="text-align: left"><u style="color:ForestGreen">despondent : glee</u></td>
      </tr>
      <tr>
        <td style="text-align: left">leisurely : haste</td>
      </tr>
      <tr>
        <td style="text-align: left">morbid : cheerfulness</td>
      </tr>
      <tr>
        <td style="text-align: left">infamous : honor</td>
      </tr>
    </tbody>
  </table>

</div>

<p>With the following logic:</p>

<div class="chatgpt-edit-block">
<div class="chatgpt-output-only">
    <blockquote>
      <p>Okay, for the pair:</p>

      <p>hospitable : courtesy</p>

      <p>The relationship is:</p>

      <p>Someone who is hospitable shows courtesy or kindness to their guests. So
there is a relationship where one concept (hospitable) leads to or results
in what the other represents (courtesy).</p>

      <p>Based on this, the choice that shows a similar relationship where one
concept results in what the other represents is:</p>

      <p>despondent : glee</p>

      <p>Because someone who is despondent lacks glee or cheer. Like hospitable leads
to courtesy, despondent precludes glee.</p>
    </blockquote>
  </div>
</div>

<p>Claude <em>correctly</em> identifies that hospitable implies showing courtesy, but
then picks the opposite relation, someone is despondent <strong>lacks</strong> glee.</p>

<p>All of Claude’s answers are <a href="/blog/sat2vec/claude_results/">here</a>.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:previous">

      <p>I used one analogy in my instruction to Claude, which explains the
discrepancy between 35 and 36. <a href="#fnref:previous" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[Word2Vec failed to solve SAT analogies, can modern language models do better? A small test of Anthropic's Claude LLM.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/sat2vec/00225-2672697451-impressionistic_painting,_four_men_studying_at_a_desk,_smoking,_looking_over_papers,_window_in_the_background.png" />
        <media:content medium="image" url="https://alexgude.com/files/sat2vec/00225-2672697451-impressionistic_painting,_four_men_studying_at_a_desk,_smoking,_looking_over_papers,_window_in_the_background.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">When Are Large Language Models Useful?</title>
      <link href="https://alexgude.com/blog/good-uses-for-large-language-models/" rel="alternate" type="text/html" title="When Are Large Language Models Useful?" />
      <published>2023-04-12T00:00:00-07:00</published>
      <updated>2023-04-12T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/good_uses_for_large_language_models</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/good-uses-for-large-language-models/"><![CDATA[<p>Large language models (LLMs) like <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a>, <a href="https://en.wikipedia.org/wiki/GPT-4#Microsoft_Bing">Bing Chat</a>, and
<a href="https://en.wikipedia.org/wiki/LaMDA">Bard</a> have gained tremendous popularity in recent months. It feels
like a pivotal moment in the technology’s growth as it becomes increasingly
integrated into people’s workflows. But despite the excitement, some people
are already dismissing the technology after they asked it questions and
received nonsensical responses. I think they are mistaken. LLMs are incredibly
valuable tools, <em>if</em> you know when to use them.</p>

<p>I gave an example in my last post of a good application for LLMs: <a href="/blog/how-i-write-with-llms-revised/">editing
prose</a>. But what specifically makes this problem ideal for solving
with a model? Succinctly, it is a problem <strong>where solving it is hard, but
verifying the solution is easy</strong>. I will go into more detail in the rest of
this post.</p>

<h2 id="what-are-they-good-for">What Are They Good For?</h2>

<p>In math, there are types of problems where finding a solution is difficult
or impossible, but confirming a solution is easy. A common strategy to solve
these problems is to guess the solution’s form and then verify it, such as for
an integral where the solution can be checked by taking its derivative.</p>

<p>Large language models are particularly useful for exactly these types of
tasks: <strong>where generating a solution is hard, but verifying it is easy</strong>.
Editing a paragraph is a prime example of this kind of task since writing
multiple versions is time-consuming, whereas verifying the quality of a single
paragraph can be done quickly.</p>

<p>Other good use cases include <a href="/blog/large-language-models-for-data-cleaning/">automating data cleaning
tasks</a> or <a href="/blog/how-i-write-code-with-llms/">writing code</a>, especially if you
have tests in place to verify the code’s correctness.</p>

<h2 id="what-are-they-bad-for">What Are They Bad For?</h2>

<p>LLMs are <strong>bad for problems where verification is hard</strong> compared to the
generation of an answer.</p>

<p>Some people are using LLMs as a replacement for search engines.<sup style="anchor-name:--fnref-mea_culpa" id="fnref:mea_culpa"><a href="#fn:mea_culpa" class="footnote" rel="footnote" role="doc-noteref">1</a></sup>
This is a perfect example of a <strong>bad use</strong> of the technology because verifying
the accuracy of the information provided by the model takes time and effort.
In fact, it often involves additional searches to confirm the validity of the
answer, which defeats the purpose of using an LLM in the first place.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:mea_culpa">

      <p><strong>Author’s Note</strong>: This specific example didn’t age well! When I wrote
this in 2023, models relied on “frozen” training weights to mimic search,
making them prone to confident hallucinations. But now, almost every model
uses live search results to ground their responses. LLMs have become a
fantastic way to search because they are now doing what they’re best at:
reading a lot more text than a human ever could and distilling it down to
the key points. <a href="#fnref:mea_culpa" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[Large language models (LLMs) are incredibly valuable tools, but they're not for everything. Here's a simple rule to know when to use them and when to avoid them.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/chatgpt/00259-1343806484-A_drawing_of_a_cute_robot_color_writing_with_a_pen_sitting_at_a_desk.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/chatgpt/00259-1343806484-A_drawing_of_a_cute_robot_color_writing_with_a_pen_sitting_at_a_desk.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">How I Write with ChatGPT</title>
      <link href="https://alexgude.com/blog/how-i-write-with-chatgpt/" rel="alternate" type="text/html" title="How I Write with ChatGPT" />
      <published>2023-02-13T00:00:00-08:00</published>
      <updated>2023-02-13T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/how_i_write_with_chatgpt</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/how-i-write-with-chatgpt/"><![CDATA[<p><strong>Note: This article describes my early experiences using ChatGPT for writing.
Since then, my methods have evolved significantly with improvements in LLMs.
For my updated approach, check out my latest post:</strong></p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/how-i-write-with-llms-revised/">
      <img src="https://alexgude.com/files/chatgpt/202502001-robot.jpg" alt="A colorful watercolor illustration on a white background of a robot sitting at a desk writing with a pen. He has a desklamp and a cup with pens in it.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/how-i-write-with-llms-revised/">
      <strong>How I Write with Large Language Models</strong>
    </a>
<br />
OpenAI's ChatGPT 3.5 transformed my writing process when it came out. After years of experience using it, I've further refined my method of using LLMs. This post explains how.  </div>
</li>
</ul>

<hr />

<p><a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a>, the new chatbot from <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a>, has received
widespread attention for its versatility in conversation, information
retrieval, and even song writing. However, its tendency to generate false but
believable information has led many to see it as <em>entertaining</em> but not
<em>useful</em>.<sup style="anchor-name:--fnref-edit_1" id="fnref:edit_1"><a href="#fn:edit_1" class="footnote" rel="footnote" role="doc-noteref">1</a></sup><sup> (Click to see my edit prompts.)</sup></p>

<p>But ChatGPT is an <strong>incredibly useful tool</strong> for my writing due to its ability
to rapidly iterate through alternative phrasings for ideas, sentences, and
paragraphs. In these cases, the model’s habit of generating false information
is not a concern because I wrote the original input so I know what it should
say, and I can easily edit out any inaccuracies introduced by the
model.<sup style="anchor-name:--fnref-edit_2" id="fnref:edit_2"><a href="#fn:edit_2" class="footnote" rel="footnote" role="doc-noteref">2</a></sup><sup style="anchor-name:--fnref-edit_3" id="fnref:edit_3"><a href="#fn:edit_3" class="footnote" rel="footnote" role="doc-noteref">3</a></sup></p>

<p>Let me illustrate how I use ChatGPT to write and copyedit my work.<sup style="anchor-name:--fnref-edit_4" id="fnref:edit_4"><a href="#fn:edit_4" class="footnote" rel="footnote" role="doc-noteref">4</a></sup></p>

<h2 id="my-writing-process">My Writing Process</h2>

<p>I start by writing a rough draft which may consist of a few bullet points or
a detailed set of paragraphs. Next, I give it to ChatGPT for refinement.</p>

<div class="chatgpt-edit-block">
<div class="chatgpt-prompt">
<strong>Prompt</strong>
<div class="chatgpt-prompt-only">
      <blockquote>
        <p>Give me three versions of this paragraph:</p>

        <p>I start by writing a rough draft. This draft could be a just a few bullet
points or completely fleshed out paragraphs. I then feed it into ChatGPT:</p>
      </blockquote>
    </div>
</div>

<div class="chatgpt-output">
<strong>Output</strong>
<div class="chatgpt-output-only">
      <blockquote>
        <ol>
          <li>I begin my writing process with a rough draft, which may range from a
simple list of bullet points to fully developed paragraphs. Next, I input
it into ChatGPT.</li>
          <li>My writing method starts with a rough draft - it could be a brief outline
or a detailed set of paragraphs. Afterwards, I pass it to ChatGPT.</li>
          <li>To start, I create a preliminary draft - it may consist of a handful of
bullet points or fully written paragraphs. Afterwards, I utilize ChatGPT
to refine it further.</li>
        </ol>
      </blockquote>
    </div>
</div>
</div>

<p>I don’t usually take the entire output from ChatGPT, rather I use it to
inspire my own editing:<sup style="anchor-name:--fnref-edit_5" id="fnref:edit_5"><a href="#fn:edit_5" class="footnote" rel="footnote" role="doc-noteref">5</a></sup></p>

<blockquote>
  <p>I start by writing a rough draft which may consist of a few bullet points
or a detailed set of paragraphs. Next, I give it to ChatGPT for refinement.</p>
</blockquote>

<p>I iterate using this process, refining my text by sending the sections I’m
unhappy with to ChatGPT for improvement until I’m satisfied.<sup style="anchor-name:--fnref-edit_6" id="fnref:edit_6"><a href="#fn:edit_6" class="footnote" rel="footnote" role="doc-noteref">6</a></sup></p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:edit_1">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>ChatGPT is the new large language model from OpenAI. It has garnered a lot of attention for its wide range of abilities from conversation to writing songs to information retrieval. It has also drawn criticism for its habit of making up plausible but completely wrong facts, leading many to conclude that while it is entertaining it is not useful.</blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. OpenAI's ChatGPT is a highly advanced language model that has received widespread attention for its diverse abilities, including conversation, songwriting, and information retrieval. However, its tendency to generate false information that appears credible has led to criticism and the conclusion that, although entertaining, it may not be practical.<br /><br />2. ChatGPT, the new AI model from OpenAI, has captured considerable attention for its wide range of skills, such as conversation, songwriting, and information retrieval. Despite this, its propensity for producing believable but false facts has sparked criticism, and many view it as entertaining but not practical.<br /><br />3. OpenAI's ChatGPT, a new and sophisticated language model, has drawn a lot of attention for its versatility in conversation, songwriting, and information retrieval. Despite this, it has also faced criticism for generating seemingly credible but completely false information, leading many to see it as entertaining but lacking practicality.</blockquote></div></div></div>
      <p><a href="#fnref:edit_1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:edit_2">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>I find ChatGPT to be incredibly useful as part of my writing process, because ChatGPT makes it fast and easy to generate different possible ways to write down an idea, sentence, or paragraph. In these cases the fact that the model makes things up is irrelevant because I wrote the original source material so I can filter out any inaccuracies the model introduces. </blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. I consider ChatGPT to be a valuable tool in my writing process due to its ability to quickly and effortlessly generate various ways of phrasing an idea, sentence, or paragraph. In these cases, the fact that the model generates false information is not a concern as I am the originator of the source material and can easily identify and remove any inaccuracies introduced by the model.<br /><br />2. For me, ChatGPT is an extremely useful aspect of my writing as it facilitates the rapid creation of alternative phrasings for ideas, sentences, or paragraphs. The model's tendency to generate false information is inconsequential in these scenarios as I am the source of the original material and can easily filter out any inaccuracies introduced by the model.<br /><br />3. ChatGPT is an indispensable part of my writing process as it enables me to quickly generate various ways to express an idea, sentence, or paragraph. Its tendency to produce false information is not a problem as I am the writer of the original source material, and I can easily eliminate any inaccuracies introduced by the model.</blockquote></div></div></div>
      <p><a href="#fnref:edit_2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:edit_3">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>I have used 'tendency' too much already. Rewrite this five times, each time only replacing the word 'tendency' with another few words to maintain the meaning:<br /><br />In these cases, the model's tendency to generate false information is not a problem because I wrote the original input so I know what it should say, and I can easily edit out any inaccuracies introduced by the model.</blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. In these cases, the model's habit of producing false information does not pose an issue because I created the original input and can easily remove any inaccuracies it may introduce.<br /><br />2. In these scenarios, the model's inclination to generate false information is of no concern as I am the writer of the original material and can effectively edit out any inaccuracies introduced by the model.<br /><br />3. The model's propensity for generating false information is not a concern in these instances as I am aware of the intended content of the original input and can efficiently remove any inaccuracies it may introduce.<br /><br />4. The model's tendency to produce false information is not a hindrance in these situations as I am familiar with the original input and can effortlessly edit out any inaccuracies it generates.<br /><br />5. The model's habit of creating false information is not problematic because I wrote the original input and can easily eliminate any inaccuracies the model introduces.</blockquote></div></div></div>
      <p><a href="#fnref:edit_3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:edit_4">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>Five versions of:<br /><br />Let me show you how I use it to write and edit my work.</blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. Allow me to demonstrate how I utilize it in writing and revising my pieces.<br /><br />2. Let me illustrate how I employ it in the creation and modification of my writing.<br /><br />3. I will show you how I use ChatGPT for writing and editing my writing projects.<br /><br />4. Let me reveal how I integrate it into my writing and editing process.<br /><br />5. I will demonstrate to you the way I utilize it for writing and revising my work.</blockquote></div></div></div>
      <p><a href="#fnref:edit_4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:edit_5">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>Three versions of:<br /><br />I generally don't take the full paragraph, but I use ChatGPT's versions to get ideas for my own:</blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. I don't usually use the entire output from ChatGPT, but I do use it as inspiration for my own writing.<br /><br />2. Typically, I don't adopt ChatGPT's output in its entirety, instead I draw ideas from it for my writing.<br /><br />3. My approach is not to fully adopt the ChatGPT output, but rather use it to spark ideas for my own writing.</blockquote></div></div></div>
      <p><a href="#fnref:edit_5" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:edit_6">
      <div class="chatgpt-edit-block"><div class="chatgpt-prompt"><strong>Prompt</strong><div class="chatgpt-prompt-only"><blockquote>Give me three versions of this:<br /><br /> I iterate on this process, taking parts of my text I'm not happy with and sending them to ChatGPT, until I like how the work flows.</blockquote></div></div><div class="chatgpt-output"><strong>Output</strong><div class="chatgpt-output-only"><blockquote>1. Revising my text with ChatGPT's help, I improve it piece by piece until I'm satisfied with the flow.<br /><br />2. With the aid of ChatGPT, I repeat the process of refining my text by sending unsatisfactory portions for improvement until I achieve the desired flow.<br /><br />3. I use ChatGPT to continuously refine my text, sending sections I'm unhappy with for improvement until I'm content with the overall flow.</blockquote></div></div></div>
      <p><a href="#fnref:edit_6" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="large-language-models" />
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[OpenAI's ChatGPT is viewed as entertaining but not useful because it makes up facts. But I find it incredibly valuable for writing. Here is how I use it.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/chatgpt/00137-2463433472-watercolor_illustration_adorable_robot_desk_lamp_sitting_at_a_typewriter_chair_desk_clear_straight_lines.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/chatgpt/00137-2463433472-watercolor_illustration_adorable_robot_desk_lamp_sitting_at_a_typewriter_chair_desk_clear_straight_lines.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">AI, Artists, and Technology</title>
      <link href="https://alexgude.com/blog/ai-artists-and-technology/" rel="alternate" type="text/html" title="AI, Artists, and Technology" />
      <published>2023-01-30T00:00:00-08:00</published>
      <updated>2023-01-30T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/ai_artists_and_technology</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/ai-artists-and-technology/"><![CDATA[<p>The open-source release of <a href="https://en.wikipedia.org/wiki/Stable_Diffusion">Stable Diffusion</a> has sparked an explosion of
progress in <a href="https://en.wikipedia.org/wiki/Artificial_intelligence_art">AI-generated art</a>. Although it is in its infancy, I can
already tell this new tool is going to revolutionize visual art creation. But
not everyone views AI art in a positive light. Many artists feel that <a href="https://twitter.com/Artofinca/status/1599730391698485248">AI art
stole their work</a><sup style="anchor-name:--fnref-stolen_quote" id="fnref:stolen_quote"><a href="#fn:stolen_quote" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and have <a href="https://arstechnica.com/information-technology/2022/12/artstation-artists-stage-mass-protest-against-ai-generated-artwork/">organized protests</a>
on popular sites like <em>ArtStation</em>. Other artists <a href="https://www.vice.com/en/article/ake9me/artists-are-revolt-against-ai-art-on-artstation">claim that AI-generated art
can’t be art</a><sup style="anchor-name:--fnref-not_art_quote" id="fnref:not_art_quote"><a href="#fn:not_art_quote" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> because it isn’t
human.</p>

<h2 id="ai-and-photography-as-art">AI and Photography as Art</h2>

<p>I come down on the side of AI-artists.</p>

<p>This is probably unsurprising because I am a machine learning engineer, it is
my job to build the types of systems these artists are using. But what is less
obvious is that my support is also because I am an artist, specifically a
landscape photographer.</p>

<p>Photography—just like AI-generated art—has <a href="https://daily.jstor.org/when-photography-was-not-art/">a complicated history as
“art”</a>. Although the first photograph was taken in 1826, it wasn’t
until 1924 that an American museum recognized the medium as art by <a href="https://en.wikipedia.org/wiki/Alfred_Stieglitz">including
photographs in its permanent collection</a>. At first artists feared
photography would replace traditional visual arts due to the ease of taking a
picture. But eventually they realized it was a useful tool that could be
combined with other art forms, even if they did not recognize photography as
an art in its own right.<sup style="anchor-name:--fnref-brush_and_pencil" id="fnref:brush_and_pencil"><a href="#fn:brush_and_pencil" class="footnote" rel="footnote" role="doc-noteref">3</a></sup><sup>, </sup><sup style="anchor-name:--fnref-the_new_path" id="fnref:the_new_path"><a href="#fn:the_new_path" class="footnote" rel="footnote" role="doc-noteref">4</a></sup></p>

<p>The concerns and criticisms currently being directed towards AI-generated art
are the same as those leveled against photography in the past. And just as
photography eventually gained acceptance as a valid form of art so will
AI-generated art. The resistance against it may be strong, but ultimately, it
is a losing battle.</p>

<h2 id="my-familys-art">My Family’s Art</h2>

<p>My family has a long history of painting. My great-great-great grandfather was
the Norwegian landscape painter <a href="https://en.wikipedia.org/wiki/Hans_Gude">Hans Gude</a>. My father, also named
<a href="https://hans-f-gude-art.github.io/website/about">Hans Gude</a>, was <a href="https://hans-f-gude-art.github.io/website/">an accomplished oil
painter</a>.<sup style="anchor-name:--fnref-hans_art" id="fnref:hans_art"><a href="#fn:hans_art" class="footnote" rel="footnote" role="doc-noteref">5</a></sup> I too wanted to make art, but I did not have
their skill with a brush so I picked up a camera instead.</p>

<p>I was drawn to photography <strong>specifically because</strong> it used technology. I like
learning new technologies and how to master them. I <em>also</em> thought it would be
easier to make art I was happy with using a camera. I have since learned that
photography has its own set of skills to master, but after 15 years I think I
was mostly right: it is much easier than oil painting.</p>

<p>I wonder what my great grandfather would think of my art. He spent months or
years creating his seascapes, while my photographs are captured in a fraction
of a second with the push of a button, and maybe a few hours adjusting tone
curve and highlights back at my computer.</p>

<p>But I like to think that he would view my work as a continuation of our
family’s artistic tradition. Maybe in the future, my descendants will find the
camera too complicated and instead compose prompts for AI to translate into
images. To me, that’s simply another evolution of the art form.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:stolen_quote">

      <p><figure class="cited-quote"><blockquote cite="https://twitter.com/Artofinca/status/1599730391698485248"> <p>Current AI “art” is created on the backs of hundreds of thousands of artists and photographers who made billions of images and spend time, love and dedication to have their work soullessly stolen and used by selfish people for profit without the slightest concept of ethics.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Nanitchkov, Alexander (@Artofinca). <a href="https://twitter.com/Artofinca/status/1599730391698485248">“Tweet”</a> <cite>Twitter</cite>. 2022-12-05.</span></figcaption></figure> <a href="#fnref:stolen_quote" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:not_art_quote">

      <p><figure class="cited-quote"><blockquote cite="https://www.vice.com/en/article/artists-are-revolt-against-ai-art-on-artstation/"> <p>“I believe art is something inherently and intrinsically human, even corporate art made-for-hire is meticulously crafted by experts in their fields,” [Nicholas] Kole said. “When we sit down to draw, design, sculpt or paint, each mark is made with an intention. Each step of the process is an opportunity to ask new questions, tune the piece to the precise context it’s intended for, to add expressiveness and even a point of view. The result—movies, shows, games—are intended to connect that intricate craft with an audience who appreciates and enjoys it.” </p>  <p>AI does none of this, he explained, and he sees “a world filling up with meaningless, regurgitative cardboard cutouts that remind us of real art.”</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Xiang, Chloe. <a href="https://www.vice.com/en/article/artists-are-revolt-against-ai-art-on-artstation/">“Artists Are Revolting Against AI Art on ArtStation”</a> <cite>Vice</cite>. 2022-12-14.</span></figcaption></figure> <a href="#fnref:not_art_quote" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:brush_and_pencil">

      <p><figure class="cited-quote"><blockquote cite="https://doi.org/10.2307/25505621"> <p>The fear has sometimes been expressed that photography would in time entirely supersede the art of painting. Some people seem to think that when the process of taking photographs in colors has been perfected and made common enough, the painter will have nothing more to do. We need not fear anything of the kind. Perfection in photography may rid us in time of all the poor work done in color. The work of the artist, however, in which is seen his own individuality, his own perception of the beautiful, his own creation in fact, can no more perish than the soul which inspired it.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Clopath, Henrietta. <a href="https://doi.org/10.2307/25505621">“Genuine Art versus Mechanism”</a> <cite>Brush and Pencil</cite>. vol. 7, no. 6. 1901-03-01. pp. 331–333. doi: <a href="https://doi.org/10.2307/25505621">10.2307/25505621</a>.</span></figcaption></figure> <a href="#fnref:brush_and_pencil" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:the_new_path">

      <p><figure class="cited-quote"><blockquote cite="https://www.jstor.org/stable/20542505"> <p>Photography is an infinitely valuable mechanism by which to obtain records of limited abstract truth, and as such, may be of great service to the artist. Much may be learned about drawing by reference to a good photograph, that even a man of quick natural perception would be slow to learn without such help. But, unless the real shortcomings of the photograph are understood, it must certainly mislead if followed.</p>  <p>But beyond these merely technical matters, art differs from any mechanical process in being “the expression of man’s delight in God’s work”, and thus it appeals to, and awakens all noble sympathy and right feeling. All labor of love must have something beyond mere mechanism at the bottom of it.</p>  </blockquote><figcaption>—<span markdown="0" class="citation"><a href="https://www.jstor.org/stable/20542505">“Art and Photography”</a> <cite>The New Path</cite>. vol. 2, no. 12. 1865-12-01. pp. 198–199.</span></figcaption></figure> <a href="#fnref:the_new_path" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:hans_art">

      <p>My father somewhat rejected the title of “artist”, although in later life
he branded himself as such. He preferred to think of himself as a
craftsman, honing his skills through hard work and study. <a href="#fnref:hans_art" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="generative-ai" />
        
          <category term="machine-learning" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[AI generated art took off with the open-source release of Stable Diffusion, leaving some artists worried. As an artist and machine learning engineer, here is my take.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/ai_artists_and_technology/field_of_yellow_by_alex_gude_painted_by_chatgpt_20250613_hans_gude.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/ai_artists_and_technology/field_of_yellow_by_alex_gude_painted_by_chatgpt_20250613_hans_gude.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: Pedestrian Safety on Halloween</title>
      <link href="https://alexgude.com/blog/switrs-pedestrian-incidents-on-halloween/" rel="alternate" type="text/html" title="SWITRS: Pedestrian Safety on Halloween" />
      <published>2022-12-01T00:00:00-08:00</published>
      <updated>2022-12-01T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_pedestrian_incidents_on_halloween</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-pedestrian-incidents-on-halloween/"><![CDATA[<p>In my last post, I found that Halloween is the <a href="/blog/switrs-pedestrian-incidents-by-date/">most dangerous day of the year
for pedestrians</a>, with a higher number of incidents than any other
day, according to <a href="/blog/switrs-sqlite-hosted-dataset/">data from SWITRS</a>. I also found that the risk of
pedestrian incidents is higher during commute hours, regardless of the date.
In this article, I will explore these patterns in more detail using the same
SWITRS data, but with a focus on Halloween.</p>

<p>As per usual, the Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-pedestrian-halloween/SWITRS%20Incident%20Dates%20With%20Pedestrians%20On%20Halloween.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-pedestrian-halloween/SWITRS%20Incident%20Dates%20With%20Pedestrians%20On%20Halloween.ipynb">rendered on Github</a>).</p>

<h2 id="data-selection">Data Selection</h2>

<p>I selected crashes involving pedestrians from the <a href="https://github.com/agude/SWITRS-to-SQLite">SQLite database</a> with
the following query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">collision_date</span><span class="p">,</span>
       <span class="n">collision_time</span><span class="p">,</span>
       <span class="n">pedestrian_killed_count</span>
<span class="k">FROM</span> <span class="n">collisions</span>
<span class="k">WHERE</span> <span class="n">Collision_Date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">pedestrian_Collision</span> <span class="o">=</span> <span class="mi">1</span>        <span class="c1">-- Involves a pedestrian</span>
<span class="k">AND</span> <span class="n">collision_date</span> <span class="o">&lt;=</span> <span class="s1">'2020-12-31'</span>  <span class="c1">-- 2021 is incomplete</span>
<span class="c1">-- and it happens on Halloween</span>
<span class="k">AND</span> <span class="n">strftime</span><span class="p">(</span><span class="s1">'%m-%d'</span><span class="p">,</span> <span class="n">Collision_Date</span><span class="p">)</span> <span class="o">=</span> <span class="s1">'10-31'</span>
</code></pre></div></div>

<p>This gave me 1168 data points, of which 64 involve a pedestrian fatality,
spanning the years 2001 through 2020. Incidents after 2020 are rejected
because the database dump comes from mid-2021, and so that year is incomplete.</p>

<h2 id="incidents-per-hour">Incidents Per Hour</h2>

<p><a href="https://web.archive.org/web/20201010053605/https://www.curbed.com/2019/10/25/20927701/halloween-safety-pedestrian-deaths-kids">Alissa Walker</a> wrote<sup style="anchor-name:--fnref-aw_quote" id="fnref:aw_quote"><a href="#fn:aw_quote" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> that it is not just drivers that make
Halloween deadly, it is commuters. The best way to explore this point is to
look at when in the day crashes happen:</p>

<p><a href="/files/switrs-pedestrian-halloween/pedestrian_incidents_by_hour_on_halloween.svg"><img src="/files/switrs-pedestrian-halloween/pedestrian_incidents_by_hour_on_halloween.svg" alt="Average number of incidents involving pedestrians per hour on Halloween
from 2001 to 2020, separated by weekend and
weekdays." /></a></p>

<p>As we saw in the data for all dates, weekdays <a href="/blog/switrs-pedestrian-incidents-by-date/#hour-by-hour">have two major peaks in
collisions during the morning and evening commutes, as well as a peak during
school pickup times</a>. Examining the data for Halloween
specifically, we see that when it falls on a weekday the three expected peaks
(morning and evening commutes, and school pick-up) are present, but there is
also a fourth peak at 18:00, likely due to a combination of darkness making it
difficult for drivers to see pedestrians and trick-or-treating bringing more
people out walking. This data supports Walker’s observation that commuter
traffic contributes significantly to the number of pedestrian incidents.</p>

<h2 id="fatality-rates">Fatality Rates</h2>

<p>But Walker makes a very specific claim: that fatalities involving children
increase on weekday Halloweens. Does the data support this claim? To find out,
we need to look at the fatality rate instead of the total number of fatalities
because the number of people driving and walking changes year-by-year and
using the rate helps to normalize some of this variation. Below is a plot of
the fatality rates for each year’s Halloween, separated into weekday and
weekend:</p>

<p><a href="/files/switrs-pedestrian-halloween/pedestrian_fatality_rate_by_day_type_on_halloween.svg"><img src="/files/switrs-pedestrian-halloween/pedestrian_fatality_rate_by_day_type_on_halloween.svg" alt="Fatality rate for pedestrians per year on Halloween separated by weekday vs
weekend." /></a></p>

<p>The data above includes all pedestrian fatalities, not just those involving
children. At first glance, the distributions for weekday and weekend Halloween
fatalities appear similar. A <a href="https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney_U_test">Mann–Whitney U test</a> confirms this, with
a <em>p</em>-value of 0.93, indicating that the difference between the two is not
statistically significant.</p>

<p>But what about children alone (defined as pedestrians under 18)? Here is that
data:</p>

<p><a href="/files/switrs-pedestrian-halloween/children_pedestrian_fatality_rate_by_day_type_on_halloween.svg"><img src="/files/switrs-pedestrian-halloween/children_pedestrian_fatality_rate_by_day_type_on_halloween.svg" alt="Fatality rate for child pedestrians per year on Halloween separated by
weekday vs weekend." /></a></p>

<p>One interesting observation is that no children have been killed by cars on
weekend Halloweens, whereas about half of the weekdays have seen at least one
child death. This suggests that there is something about weekday Halloweens
that makes them particularly dangerous for children, consistent with Walker’s
claim.</p>

<p>Despite this, the data does not show a significant difference between the two
distributions, with a <em>p</em>-value of 0.08. However, this lower <em>p</em>-value as
compared to the all-ages data does indicate some evidence for the specific
claim about child deaths.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:aw_quote">

      <p><figure class="cited-quote"><blockquote cite="https://web.archive.org/web/20201010053605/https://www.curbed.com/2019/10/25/20927701/halloween-safety-pedestrian-deaths-kids"> <p>But when the commuting drivers are removed from the equation, deaths seem to go down. A study by AutoInsurance.org used FARS data to compare 24 years of crash data by days of the week. Halloweens that fell on workdays had an 83 percent increase in deadly crashes involving kids compared to weekend days. The worst day? Friday. Since 1994, the three deadliest Halloween nights for kids have all been Friday nights.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Walker, Alissa. <a href="https://web.archive.org/web/20201010053605/https://www.curbed.com/2019/10/25/20927701/halloween-safety-pedestrian-deaths-kids">“The most terrifying part of Halloween for kids is our deadly streets”</a> <cite>Curbed</cite>. Vox Media. October 25, 2019.</span></figcaption></figure> <a href="#fnref:aw_quote" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Halloween can be a dangerous time for pedestrians. In this post, I explore the statistics on pedestrian-vehicle collisions, including when these incidents are most likely to occur.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-pedestrian-halloween/auto_accident_loc_2016819574_1920.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-pedestrian-halloween/auto_accident_loc_2016819574_1920.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: On What Days Do Drivers Hit Pedestrians?</title>
      <link href="https://alexgude.com/blog/switrs-pedestrian-incidents-by-date/" rel="alternate" type="text/html" title="SWITRS: On What Days Do Drivers Hit Pedestrians?" />
      <published>2022-11-10T00:00:00-08:00</published>
      <updated>2022-11-10T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_pedestrian_incidents_by_date</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-pedestrian-incidents-by-date/"><![CDATA[<p>It has been a while since I have used the <a href="/blog/switrs-sqlite-hosted-dataset/">SWITRS data</a> to look at
vehicle collisions in California. Since my last article—<a href="/blog/switrs-bicycle-crashes-by-date/"><em>On What Days Do
Cyclists Crash?</em></a>—we have lived through a massive
<a href="https://en.wikipedia.org/wiki/COVID-19_pandemic_in_California">pandemic</a> that <em>significantly changed</em> how people drive, including <a href="/blog/switrs-covid-19-lockdown-fatal-traffic-collisions/">a
huge increase to fatalities</a> and changing <a href="/blog/switrs-ford-vs-toyota-during-covid-19/">what type of
drivers have crashes</a>.</p>

<p>So with all the new data, I wanted to look at the most vulnerable road users:
pedestrians.</p>

<p>As per usual, the Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-pedestrian-incidents-by-date/SWITRS%20Incident%20Dates%20With%20Pedestrians.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-pedestrian-incidents-by-date/SWITRS%20Incident%20Dates%20With%20Pedestrians.ipynb">rendered on Github</a>).</p>

<h2 id="data-selection">Data Selection</h2>

<p>I selected crashes involving pedestrians from the <a href="https://github.com/agude/SWITRS-to-SQLite">SQLite database</a> with
the following query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">collision_date</span>
     <span class="p">,</span> <span class="n">collision_time</span>
     <span class="p">,</span> <span class="n">pedestrian_killed_count</span>
<span class="k">FROM</span> <span class="n">collisions</span>
<span class="k">WHERE</span> <span class="n">collision_date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">pedestrian_Collision</span> <span class="o">=</span> <span class="mi">1</span>        <span class="c1">-- Involves a pedestrian</span>
<span class="k">AND</span> <span class="n">collision_date</span> <span class="o">&lt;=</span> <span class="s1">'2020-12-31'</span>  <span class="c1">-- 2021 is incomplete</span>
</code></pre></div></div>

<p>This gave me 282,039 data points to examine spanning the years 2001 through
2020. Incidents after 2020 are rejected because the database dump comes from
mid-2021, and so that year is incomplete.</p>

<h2 id="crashes-per-week">Crashes per Week</h2>

<p>For bicycle involved incidents, <a href="/blog/switrs-bicycle-crashes-by-date/#crashes-per-week">I found there was an increase from about 2008
through 2013 followed by a decrease</a>. For both bicycles <a href="/blog/switrs-motorcycle-crashes-by-date/#crashes-per-week">as well as
motorcycles</a>, I found strong seasonality with many more crashes during
the summer when people are out riding to take advantage of the weather.
Pedestrian involved incidents defy both these trends:</p>

<p><a href="/files/switrs-pedestrian-incidents-by-date/pedestrian_incidents_per_month_in_california.svg"><img src="/files/switrs-pedestrian-incidents-by-date/pedestrian_incidents_per_month_in_california.svg" alt="Step plot showing pedestrian involved incidents per week in California from
2001 through 2020." /></a></p>

<p>Pedestrian involved incidents were flat or slightly down until about 2013,
when instead of decreasing like bicycle collisions they <strong>increased
strongly</strong>. Like both bicycles and motorcycle crashes,
pedestrian incidents are strongly seasonal but they decrease in the summer
(when there is a lot of light for drivers to see pedestrians) and <strong>increase
in the winter</strong> when it gets dark early and drivers can’t see them. Of course
there is also a massive decrease when COVID restrictions kept most people home
starting in March 2020.</p>

<h2 id="day-by-day">Day-by-day</h2>

<p><a href="/blog/switrs-crashes-by-date/#day-by-day">Cars are involved in crashes</a> on days when drivers have to commute
to work <em>and</em> on holidays where people travel. The worst day is Halloween when
people work and then go out and have fun after. I was curious if the large
increase on Halloween was due to a large increase in pedestrian collisions;
the answer is no:</p>

<p><a href="/files/switrs-pedestrian-incidents-by-date/mean_pedestrian_incidents_by_date.svg"><img src="/files/switrs-pedestrian-incidents-by-date/mean_pedestrian_incidents_by_date.svg" alt="Step plot showing mean pedestrian involved incidents by date in California
average over 2001 through 2020." /></a></p>

<p>On Halloween, drivers are more likely to hit pedestrians than on any other day
of the year. But its only about 15 to 20 more incidents than on any other
October day, and there are almost 200 additional car crashes on Halloween. The
number of additional pedestrian incidents does not account for much of the
increase in car crashes. For a more detailed analysis of pedestrian collisions
on Halloween, <a href="/blog/switrs-pedestrian-incidents-on-halloween/">check out my other post</a>.</p>

<p>Otherwise there are some interesting patterns. Many holidays trend the same
direction as cars: New Years, Memorial Day, Veterans Day, Thanksgiving, and
Christmas all see a large reduction in both car crashes and pedestrian
incidents. Halloween, as covered above, sees a large increase in both.</p>

<p>One outlier is the 4th of July. Car crashes decrease because people do not
have to commute, but pedestrian incidents increase. I think this is because
people are walking around in the dark going to and coming back from watching
fireworks, and drivers have trouble seeing them.</p>

<h2 id="hour-by-hour">Hour-by-hour</h2>

<p>Finally we can look at when cars hit pedestrians by hour and whether it is a
weekend or not:</p>

<p><a href="/files/switrs-pedestrian-incidents-by-date/pedestrian_incidents_by_hour.svg"><img src="/files/switrs-pedestrian-incidents-by-date/pedestrian_incidents_by_hour.svg" alt="Histogram showing the number of pedestrian involved incidents on average
per hour of the day for weekends and weekdays," /></a></p>

<p>The most striking feature is the large increase in the number of incidents
during the morning commute (07–09) and again in the during the evening
commute (17–19)! Commuters in cars are dangerous to pedestrians!</p>

<p>There is also an increase in incidents on both weekends and weekdays at about
17:00. This is probably because that is around sunset.<sup style="anchor-name:--fnref-sunset" id="fnref:sunset"><a href="#fn:sunset" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>The weekend curve rises smoothly through the day, but the weekday curve has a
large increase at 14. I suspect this is from school pickup which is generally
earlier than the commute.</p>

<p>Finally, it is interesting that the number of late night and early morning
incidents is much higher on the weekend. This is likely due to people going
out to bars, as the number drops off at 02:00 which is <a href="https://en.wikipedia.org/wiki/Last_call">when the bars close in
California</a>.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:sunset">

      <p>This dataset covers the whole year so the exact time of sunset changes. It
would be interesting to make a similar chart but relative to sunrise and
sunset. <a href="#fnref:sunset" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Being a pedestrian is dangerous in a world built for automobiles. In this post explore how pedestrian-involved collisions have trended in time. Take a look!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-pedestrian-incidents-by-date/auto_accident_loc_2016842389_1926.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-pedestrian-incidents-by-date/auto_accident_loc_2016842389_1926.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Using Scikit-learn Pipelines with Pandas Dataframes</title>
      <link href="https://alexgude.com/blog/using-sklearn-pipelines-with-pandas-dataframes/" rel="alternate" type="text/html" title="Using Scikit-learn Pipelines with Pandas Dataframes" />
      <published>2022-10-24T00:00:00-07:00</published>
      <updated>2022-10-24T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/using_sklearn_pipelines_with_pandas_dataframes</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/using-sklearn-pipelines-with-pandas-dataframes/"><![CDATA[<p><a href="https://scikit-learn.org">Scikit-learn</a> is a popular Python library for training machine
learning models. <a href="https://pandas.pydata.org/">Pandas</a> is a popular Python library for manipulating
tabular data. They work great together because when you are building a machine
learning model you start by working with the data in Pandas and then when it
is cleaned up you train the model in scikit-learn.</p>

<p>But one hard thing to do is to make sure you apply the exact same data
manipulating steps to the training set as to the test set and the live data
when the model is deployed. It is very easy to <a href="https://en.wikipedia.org/wiki/Leakage_(machine_learning)">leak data</a> or to forget
a step, either of which can ruin your model.</p>

<p>To help solve this problem, Scikit-learn developed <a href="https://scikit-learn.org/stable/modules/generated/sklearn.pipeline.Pipeline.html">Pipelines</a>.
Pipelines allow you to define a sequence of transforms, including a model
training step, that is easy to apply consistently. This post will go over how
to use Pipelines with Pandas Dataframes.</p>

<h2 id="pandas-and-pipelines-formerly-not-so-simple">Pandas and Pipelines: Formerly Not So Simple</h2>

<p>It used to be tough to use Pandas Dataframes and scikit-learn pipelines
together. There <a href="https://scikit-learn.org/stable/modules/generated/sklearn.compose.ColumnTransformer.html#sklearn.compose.ColumnTransformer">was a <code class="language-plaintext highlighter-rouge">ColumnTransformer</code></a> to work with
dataframes, but it had some major limitations since the output of the
transformer was a numpy array. This meant that if you used a second
<code class="language-plaintext highlighter-rouge">ColumnTransformer</code> in your pipeline you would get the following error:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>ValueError: Specifying the columns using strings is
only supported for pandas DataFrames
</code></pre></div></div>

<p>But scikit-learn version 1.2 <a href="https://github.com/scikit-learn/scikit-learn/pull/23734">updated the pipeline API</a> to fix this! Now
there is the option to output Pandas dataframes!</p>

<h2 id="a-working-pipeline">A working pipeline</h2>

<p>Now that the <a href="https://scikit-learn.org/dev/auto_examples/miscellaneous/plot_set_output.html"><code class="language-plaintext highlighter-rouge">set_output</code> API</a> exists, we can chain
<code class="language-plaintext highlighter-rouge">ColumnTransformer</code> without error!</p>

<p>For example, we can impute for one column, and then scale it and a few others.
First we set up the two <code class="language-plaintext highlighter-rouge">ColumnTransformer</code>, one to impute and one to scale:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Apply each feature pipeline using a column transform
</span><span class="n">imputer</span> <span class="o">=</span> <span class="p">(</span>
  <span class="sh">"</span><span class="s">imputer</span><span class="sh">"</span><span class="p">,</span>
  <span class="nc">ColumnTransformer</span><span class="p">(</span>
    <span class="p">[(</span><span class="sh">"</span><span class="s">col_impute</span><span class="sh">"</span><span class="p">,</span> <span class="nc">SimpleImputer</span><span class="p">(),</span> <span class="p">[</span><span class="sh">"</span><span class="s">x1</span><span class="sh">"</span><span class="p">])],</span>
    <span class="n">remainder</span><span class="o">=</span><span class="sh">"</span><span class="s">passthrough</span><span class="sh">"</span><span class="p">,</span>
  <span class="p">),</span>
<span class="p">)</span>

<span class="n">scaler</span> <span class="o">=</span> <span class="p">(</span>
  <span class="sh">"</span><span class="s">scaler</span><span class="sh">"</span><span class="p">,</span>
  <span class="nc">ColumnTransformer</span><span class="p">(</span>
    <span class="p">[</span>
      <span class="p">(</span>
        <span class="sh">"</span><span class="s">col_scale</span><span class="sh">"</span><span class="p">,</span>
        <span class="nc">StandardScaler</span><span class="p">(),</span>
        <span class="p">[</span><span class="sh">"</span><span class="s">col_impute__x1</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">remainder__x2</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">remainder__x3</span><span class="sh">"</span><span class="p">],</span>
      <span class="p">)</span>
    <span class="p">],</span>
    <span class="n">remainder</span><span class="o">=</span><span class="sh">"</span><span class="s">passthrough</span><span class="sh">"</span><span class="p">,</span>
  <span class="p">),</span>
<span class="p">)</span>
</code></pre></div></div>

<p>Then we combined them in a pipeline:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">pipe</span> <span class="o">=</span> <span class="nc">Pipeline</span><span class="p">(</span>
    <span class="n">steps</span><span class="o">=</span><span class="p">[</span>
        <span class="n">imputer</span><span class="p">,</span>
        <span class="n">scaler</span><span class="p">,</span>
    <span class="p">]</span>
<span class="p">).</span><span class="nf">set_output</span><span class="p">(</span><span class="n">transform</span><span class="o">=</span><span class="sh">"</span><span class="s">pandas</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>And it works! There are two tricks; we have to:</p>

<ol>
  <li>
    <p>Make the output of each step a dataframe with <code class="language-plaintext highlighter-rouge">set_output(transform="pandas")</code>.</p>
  </li>
  <li>
    <p>Adjust the columns names of the downstream steps because they get prepended
with the name of the previous steps they’ve gone through.</p>
  </li>
</ol>

<h2 id="complete-example">Complete Example</h2>

<p>Here is a <a href="/files/pandas-pipelines//pandas_pipeline_example.ipynb">Jupyter notebook</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/pandas-pipelines//pandas_pipeline_example.ipynb">rendered on Github</a>)
with a toy dataset and a full Pandas pipeline example. Hope it helps!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="software-development" />
        
          <category term="machine-learning" />
        
          <category term="machine-learning-engineering" />
        
      

      

      
      
        <summary type="html"><![CDATA[Pandas and scikit-learn are two important libraries for building machine learning models. Here is how to get them to work together.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/pandas-pipelines/navy_pipes.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/pandas-pipelines/navy_pipes.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Computing Machine Learning Features in Real-time</title>
      <link href="https://alexgude.com/blog/realtime-machine-learning-features/" rel="alternate" type="text/html" title="Computing Machine Learning Features in Real-time" />
      <published>2022-09-19T00:00:00-07:00</published>
      <updated>2022-09-19T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/realtime_machine_learning_features</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/realtime-machine-learning-features/"><![CDATA[<p>Machine learning models are excellent at automating simple, high frequency
decisions, like:</p>

<ul>
  <li>
    <p>Should I allow this transaction?</p>
  </li>
  <li>
    <p>Should I allow this login?</p>
  </li>
  <li>
    <p>What items should I show this customer?</p>
  </li>
</ul>

<p>To make these decisions they need information about the event that they are
scoring. These pieces of information are called “<a href="https://en.wikipedia.org/wiki/Feature_(machine_learning)">features</a>”.</p>

<h2 id="up-to-date-features-and-real-time-features">Up-to-date Features and Real-time Features</h2>

<p>Some features are easy to update. Features like “<em>The number of times the user
has received a payment in the past week</em>” are just counts that can be
maintained in a feature store. When a new payment comes in, add one to the
value. When another day has passed, subtract all the old payments out. Simple.</p>

<p>But some features are hard to keep up-to-date. What if we needed the value of
the received payment feature above, but to make a decision on <em>a received
payment</em>. Whatever system is keeping count now has to be very fast so it can
update the feature and return the value to the model in order to make a
decision without keeping the user waiting too long. In a payment system “too
long” could be much less than a second!</p>

<p>Worse, some features can not be pre-computed at all. For example, the feature
“<em>Has this user ever logged in from this location?</em>” is very useful for
stopping account takeover fraud, but the model needs to make a decision before
the login completes. However, the feature can only be computed once the login
has started because only then do we know the location!</p>

<p>These features, that have to be computed during the event that a decision is
being made on, are called <strong>real-time features</strong>. I will talk about one way to
compute them below.</p>

<h2 id="a-machine-learning-system">A Machine Learning System</h2>

<p>Let’s consider a simplified machine learning system that looks like this:</p>

<p><a href="/files/realtime-features//batched_feature_computation.svg"><img src="/files/realtime-features//batched_feature_computation.svg" alt="A diagram showing a machine learning system with incoming events, a feature
store, and a model host." /></a></p>

<p>The <strong>events</strong> (in purple and green on the diagram) are a never ending stream.
Imagine the line of event boxes moving from right to left, each one taking a
turn to dump its data into the <strong>feature store</strong> (in orange). The feature
store uses the data in the events to update the values of features and reports
those values to other parts of the system.</p>

<p>The <strong>model host</strong> (in blue) also operates on events but it handles them as
they are generated, before they even have time to get to the feature store.
The event that the model is currently making a decision on is the <strong>target
event</strong> (in green).</p>

<p>The target event will eventually dump its data into the feature store
(represented by a dotted line on the diagram) but has not done so yet. The
target event does send some data to the model host (generally something simple
like a user ID or event ID so the model knows what features to get) but this
is fast compared to waiting for the feature store.</p>

<h3 id="machine-learning-model-host">Machine Learning Model Host</h3>

<p>Let’s take a closer look at the model host:</p>

<p><a href="/files/realtime-features//batched_model_host.svg"><img src="/files/realtime-features//batched_model_host.svg" alt="A diagram focusing on the model host component of the above diagram. It
shows how data handling code brings in features before passing them to the
model itself." /></a></p>

<p>The <strong>data handling code</strong> (yellow in the diagram) gets IDs and other data
from the target event so that it knows what features to get. For example, it
might get a user ID so it can get all the features associated with that
specific user. It then passes the features it receives from the feature store
to the <strong>machine learning model</strong> (red in the diagram), which makes a decision
and sends it back to the target event (or the system handling the target
event).</p>

<h2 id="real-time-computation">Real-time Computation</h2>

<p>To use this system to calculate real-time features, we make three changes:</p>

<ul>
  <li>
    <p>Add “proto-features” to the feature store.</p>
  </li>
  <li>
    <p>Send more data from the target event to the model host.</p>
  </li>
  <li>
    <p>Do additional processing in the data handling code to combine the
proto-features and data from the target event into a real-time feature.</p>
  </li>
</ul>

<p>The diagram changes very slightly:</p>

<p><a href="/files/realtime-features//realtime_model_host.svg"><img src="/files/realtime-features//realtime_model_host.svg" alt="A diagram focusing on the model host component of the first diagram. It
shows what modifying the system to calculate real-time features would look
like, with additional data handling code, proto-feature inputs, and more data
from the target event." /></a></p>

<p>Here is how that would work for our example login feature:</p>

<ul>
  <li>
    <p>Add a proto-feature that is a list of previous login locations.</p>
  </li>
  <li>
    <p>Add the current location to the data passed in from the target event.</p>
  </li>
  <li>
    <p>Update the data handling code to check if the current location is in the
list of previous login locations from the proto-feature.</p>
  </li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="machine-learning" />
        
          <category term="machine-learning-engineering" />
        
      

      

      
      
        <summary type="html"><![CDATA[Models often derive great value from real-time features, but computing them is hard because it has to be done quickly. Here is one way I have done it successfully.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/realtime-features/20110727-14-38-39--customs_house_tower_close.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/realtime-features/20110727-14-38-39--customs_house_tower_close.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Plotting the 2022 Tour de France</title>
      <link href="https://alexgude.com/blog/2022-tour-de-france-plot/" rel="alternate" type="text/html" title="Plotting the 2022 Tour de France" />
      <published>2022-08-22T00:00:00-07:00</published>
      <updated>2022-08-22T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/2022_tour_de_france_plot</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/2022-tour-de-france-plot/"><![CDATA[<p>The 109th edition of the <a href="https://en.wikipedia.org/wiki/2022_Tour_de_France">Tour de France</a> just wrapped up. It was the
first edition in several years that felt almost normal—the 2020 edition had
been delayed by <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">COVID</a> and 2021 edition moved to avoid the
<a href="https://en.wikipedia.org/wiki/2020_Summer_Olympics">Olympics</a>.</p>

<p><a href="/blog/2021-tour-de-france-plot/">Like last year</a>, let’s explore how the race unfolded with data.</p>

<p>The code that generated the plots can be found <a href="/files/tour-de-france//Tour%20de%20France%202022%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/tour-de-france//Tour%20de%20France%202022%20Plot.ipynb">rendered on Github</a>). The data <a href="/files/tour-de-france//2022-tdf-dataframe.json">is here</a>.</p>

<h2 id="the-race-for-yellow">The Race for Yellow</h2>

<p>The <a href="https://en.wikipedia.org/wiki/General_classification_in_the_Tour_de_France">yellow jersey</a>, awarded to the rider with the lowest combined
time across the 21 stages of the race, was won by <a href="https://en.wikipedia.org/wiki/Tadej_Poga%C4%8Dar">Tadej Pogačar</a> the
past two years, so he was the favorite going into this tour.</p>

<p>Other favorites this year were:</p>

<ul>
  <li>
    <p><a href="https://en.wikipedia.org/wiki/Primo%C5%BE_Rogli%C4%8D">Primož Roglič</a>, who came in second in 2020.</p>
  </li>
  <li>
    <p><a href="https://en.wikipedia.org/wiki/Jonas_Vingegaard">Jonas Vingegaard</a>, who came in second last year.</p>
  </li>
  <li>
    <p><a href="https://en.wikipedia.org/wiki/Geraint_Thomas">Geraint Thomas</a>, who won in 2018 and finished second in 2019.</p>
  </li>
  <li>
    <p><a href="https://en.wikipedia.org/wiki/Aleksandr_Vlasov_(cyclist)">Aleksandr Vlasov</a>, who won the Tour de Romandie earlier in the year.</p>
  </li>
  <li>
    <p><a href="https://en.wikipedia.org/wiki/Daniel_Mart%C3%ADnez_(cyclist)">Daniel Martínez</a>, who won the Tour of the Basque Country earlier
in the year.</p>
  </li>
</ul>

<p>After three weeks of racing, here is how the top five riders fared through the
stages:</p>

<p><a href="/files/tour-de-france//2022_tour_de_france_top_5.svg"><img src="/files/tour-de-france//2022_tour_de_france_top_5.svg" alt="A line plot showing how far behind the leader each top-finishing rider was
after each stage of the 2022 Tour de France." /></a></p>

<p>Pogačar started strong, taking the yellow Jersey on stage 6 after beating out
Vingegaard in the final sprint. He held it through the first few Alp stages
before getting isolated on stage 11. Roglič and Vingegaard, both riding for
team <a href="https://en.wikipedia.org/wiki/Daniel_Mart%C3%ADnez_(cyclist)">Jumbo-Visma</a>, force Pogačar to chase them, pulling him away from
his supporting teammates. The duo then threw attack after attack at the yellow
jersey, forcing him to constantly accelerate. By the time they got to the
final climb on the <a href="https://en.wikipedia.org/wiki/Col_du_Galibier">Col du Galibier</a>, Pogačar was so tired that
Vingegaard was able to drop him and gain nearly three minutes, taking the
yellow jersey.</p>

<p>Pogačar and Thomas stayed neck-and-neck until stage 16 when Vingegaard and
Pogačar attacked in the Pyrenees and dropped the other contenders.</p>

<h3 id="the-green-jersey">The Green Jersey</h3>

<p>The race for the <a href="https://en.wikipedia.org/wiki/Points_classification_in_the_Tour_de_France">green jersey</a>, awarded to the rider with the most
sprint points which are earned by winning intermediate sprints and stages, was
relatively boring this year. <a href="https://en.wikipedia.org/wiki/Wout_van_Aert">Wout van Aert</a> ran away with it, never
losing the jersey after winning it on stage 2.</p>

<p>Here is how the sprint race turned out, with sprint stages shaded in grey:</p>

<p><a href="/files/tour-de-france//2022_tour_de_france_top_5_sprint.svg"><img src="/files/tour-de-france//2022_tour_de_france_top_5_sprint.svg" alt="A line plot showing how far behind the points leader the top five sprint
sprinters were." /></a></p>

<h2 id="the-rest-of-the-race">The Rest of the Race</h2>

<p>176 riders started the race and just 136 finished, the lowest number to finish
since 2000. Here is how each rider fared:</p>

<p><a href="/files/tour-de-france//2022_tour_de_france.svg"><img src="/files/tour-de-france//2022_tour_de_france.svg" alt="A line plot showing how far behind the leader every rider was for each
stage." /></a></p>

<p>Wout van Aert, the green jersey winner, paced himself well. He finished just
above the middle of the pack. This is in contrast to last year’s winner, <a href="https://en.wikipedia.org/wiki/Mark_Cavendish">Mark
Cavendish</a>, who finished near the bottom that year, because he is a
much weaker climber than van Aert. Van Aert is an amazing talent, able to
climb and sprint!</p>

<p>Another sprinter, <a href="https://en.wikipedia.org/wiki/Caleb_Ewan">Caleb Ewan</a>, got the <a href="https://en.wikipedia.org/wiki/Lanterne_rouge">lanterne rougue</a>, a
prize awarded to the last place rider. Normally a strong contender, he crashed
hard on stage 13 and struggled to stay with the other riders, but still
managed to cross the finish line in Paris.</p>

<p>Pre-race hopeful Primož Roglič crashed and dislocated his shoulder on stage 5.
He held on for 10 days—thankfully for Vingegaard who used Roglič to weaken
Pogačar on stage 11—but dropped out during stage 15.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[The 2022 Tour de France saw a new winner, Jonas Vingegaard! See how he won in this post!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_italian_team.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_italian_team.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Comparing Zillow, Redfin, and Realtor.com Price Estimates in Time</title>
      <link href="https://alexgude.com/blog/online-realtor-estimate-timeseries-detailed/" rel="alternate" type="text/html" title="Comparing Zillow, Redfin, and Realtor.com Price Estimates in Time" />
      <published>2022-07-16T00:00:00-07:00</published>
      <updated>2022-07-16T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/online_realtor_estimate_timeseries_detailed</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/online-realtor-estimate-timeseries-detailed/"><![CDATA[<p>A few months ago I <a href="/blog/online-realtor-estimate-timeseries/">built a time series</a> of a house’s
price estimate from Zillow and Redfin. But there were some problems:</p>

<ul>
  <li>
    <p>I only collected dense data starting right before the property went pending.</p>
  </li>
  <li>
    <p>I collected data by hand so I often missed days.</p>
  </li>
</ul>

<p>When another nearby house put a sign out saying “coming soon”, I wrote a
script to automate the scraping and collected a much denser time series from
Zillow, Redfin, and Realtor.com. Let’s see what we can learn with more
complete data!</p>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Timeseries%20Plot%20Automated.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Timeseries%20Plot%20Automated.ipynb">rendered on Github</a>). The data can be found
<a href="/files/online-realtor-estimate-comparison//home_price_estimate_20220701.json">here</a>.</p>

<h2 id="data-collection">Data Collection</h2>

<p>I wrote a script to download the entire page for the specific house from each
of the three sites. I ran it on my Raspberry Pi everyday using <code class="language-plaintext highlighter-rouge">cron</code>. I
parsed the HTML using Python and a wrote the <a href="/files/online-realtor-estimate-comparison//home_price_estimate_20220701.json">cleaned data</a> to JSON.
That parsing notebook can be found <a href="/files/online-realtor-estimate-comparison//parse_zillow.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/online-realtor-estimate-comparison//parse_zillow.ipynb">rendered on
Github</a>). I won’t include the raw data, you will have to collect
some yourself.</p>

<h2 id="plot">Plot</h2>

<p>Here is a plot comparing the estimates for the sales price from Zillow and
Redfin in time:</p>

<p><a href="/files/online-realtor-estimate-comparison//home_price_estimate_timeseries_comparison_automated.svg"><img src="/files/online-realtor-estimate-comparison//home_price_estimate_timeseries_comparison_automated.svg" alt="A plot showing three time series, one is the estimated value of a house
according to Zillow, one is the estimate for the same house from
Redfin, and the last is the estimate from Realtor.com." /></a></p>

<p>The daily price estimates from Redfin are shown using red circles, the
estimates for Zillow are shown using blue triangles, the estimates for
Realtor.com are shown using purple diamonds.</p>

<h3 id="comments">Comments</h3>

<p>I wrote a script to ensure that I would get data for every day, but as you can
see there are still many missing points.</p>

<h4 id="realtorcom">Realtor.com</h4>

<p>Realtor.com is the most frustrating! As soon as the house was actually listed
on the market they stopped providing estimates! Realtor.com starts estimating
again only after the sale price is posted. <strong>This defeats the entire point!</strong>
The estimate is most important when the house is actually for sale and
Realtor.com just punts completely. Embarrassing.</p>

<p>Still we can see that, when the house is <em>not</em> for sale, they update their
estimate roughly every two weeks. Their initial estimates are not too bad,
just about 6% low from the final sale price.</p>

<h4 id="zillow">Zillow</h4>

<p>Zillow similarly is missing estimates for most of the time when the home is
actually for sale. Their page shows <em>“Zestimate: None”</em> with an error
explanation blaming county transactional data.<sup style="anchor-name:--fnref-error" id="fnref:error"><a href="#fn:error" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> I am sure that’s <em>true</em>
but I am unimpressed. Dealing with missing data is a key part of building a
robust machine learning model.</p>

<p>Zillow’s model updates more frequently than Realtor.com’s. It slowly climbs
until it abruptly stops estimating a few days after the listing is posted.
This suggests they make use of a different model for currently on-the-market
homes and that that model requires more and different data than the
off-the-market model.</p>

<p>The Zillow estimate does return at the end of the pending period but… I do
not have anything nice to say about it. Just look at that variance!</p>

<p>Zillow underestimates the final price by about 10%.</p>

<h4 id="redfin">Redfin</h4>

<p>Redfin is the only company that keeps posting estimates once the house is
actually for sale! Their pre-listing estimate is almost exactly right, but
once the house is listed their on-the-market model overestimates by about
10%.</p>

<p>Before the listing is posted, Redfin updates its estimate roughly weekly, and
like Zillow it takes 5 days to switch to the on-the-market model. This
suggests that both sites get their data indicating the house is for sale from
the same source. After the listing the model updates daily.</p>

<p>The on-the-market model trends upwards at first and then stabilizes after the
house is pending, but because the time between listing and pending was so
short it is impossible to tell if the stabilization was due to the listing
going pending or not.</p>

<h2 id="conclusions">Conclusions</h2>

<p>With the denser data and all three sites to compare, I conclude the following:</p>

<ul>
  <li>
    <p>Zillow and Redfin both use separate models for the time pre/post-listing
and the time when the house is listed. It’s likely Realtor.com does as well
which is why it stopped updating as soon as the house was listed.</p>
  </li>
  <li>
    <p>The on-the-market model, used during the listing period, are updated more
frequently, this is probably due to the fact that they have more
high-frequency data (views, time on market, etc.) and the possibility of
making a commission on the property increases the amount of money they’re
willing to spend running the models.</p>
  </li>
  <li>
    <p>Zillow and Redfin are both using the same data source to determine when to
switch their models, and it appears to be different than the source they use
to display if a home is actually listed or not. I don’t know why that would
be. Possibly their models require data that is not immediately available?</p>
  </li>
  <li>
    <p>Zillow and Realtor.com require the same data for their model during the
listing period and fail if that source is unavailable. Redfin continues
estimating, but does so poorly.</p>
  </li>
</ul>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:error">
      <p>The error message:</p>

      <blockquote>
        <p><strong>Where’s the Zestimate?</strong></p>

        <p>County transactional data for this home is insufficient so we cannot
calculate a Zestimate. We are adding data all the time, so be sure to come
back.</p>
      </blockquote>
      <p><a href="#fnref:error" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[How do various online brokers' home price estimates change in time? I use a recently sold house near my neighborhood to find out. Come check it out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_mission_style_house.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_mission_style_house.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">A Career Involves Luck: My Annotated Resume</title>
      <link href="https://alexgude.com/blog/my-luck-resume/" rel="alternate" type="text/html" title="A Career Involves Luck: My Annotated Resume" />
      <published>2022-06-13T00:00:00-07:00</published>
      <updated>2022-06-13T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/my_luck_resume</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-luck-resume/"><![CDATA[<div class="resume">

  <div class="fake-h1">Alexander Gude</div>

  <div class="subtitle">Lucky Data Scientist / Machine Learning Engineer</div>

  <h2 id="statement">Statement</h2>

  <p>I am currently a machine learning engineer at a major tech company. My career
has taken a winding path, but I am very happy with where I am and where I am
headed. I have worked hard, but that alone does not explain my success; I have
had many <strong>lucky breaks</strong>, times when some event completely outside of my
control worked out in my favor.</p>

  <p>I think it is important to remember that no one is entirely self made; every
career has some element of luck involved. I have cataloged my lucky breaks on
this page in the form of a too-honest resume.</p>

  <h2 id="education">Education</h2>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>University of California, Berkeley</h3></div>
    <div class="resume-location">Berkeley, CA</div>
    <div class="resume-position">BA, Physics (Honors), College of Letters and Sciences</div>
    <div class="resume-dates"><p>2004–2008</p></div>
</div>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>University of Minnesota</h3></div>
    <div class="resume-location">Minneapolis, MN</div>
    <div class="resume-position">PhD, High Energy Particle Physics</div>
    <div class="resume-dates"><p>2009–2015</p></div>
</div>

  <ul>
    <li>My <a href="/blog/a-career-starts-with-rejection/#to-grad-school">GRE score was pretty bad</a>. In 2008 I was rejected from every
grad school I applied to. In 2009, <a href="https://www.physics.umn.edu/people/yk.html">Yuichi Kubota</a> saw my application to
Minnesota and thought “we can give him a shot”. He did that for a lot of
people my year.</li>
  </ul>

  <h2 id="experience">Experience</h2>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>Supernova Cosmology Project</h3></div>
    <div class="resume-location">Berkeley, CA</div>
    <div class="resume-position">Undergraduate Research Assistant</div>
    <div class="resume-dates"><p>2005-2009</p></div>
</div>

  <ul>
    <li>
      <p>My friend’s aunt worked at Lawrence Berkeley Labs as an executive assistant.
My resume got passed to her (I don’t even recall how exactly) where it
eventually found its way to the <a href="https://en.wikipedia.org/wiki/Supernova_Cosmology_Project">Supernova Cosmology Project</a> under
Nobel prize winner <a href="https://en.wikipedia.org/wiki/Saul_Perlmutter">Saul Perlmutter</a>. Saul’s post doc, Nao Suzuki,
wanted a research assistant and invited me up to the lab.</p>
    </li>
    <li>
      <p>Weirdly, and lucky for me, Nao had decided not to use <a href="https://en.wikipedia.org/wiki/IDL_(programming_language)">IDL</a>, the
standard language for astrophysics software, and instead wanted to work with
Python. We used Numarray and Numeric (which would later become Numpy),
giving me a head start in the skills I’d need later for machine learning.</p>
    </li>
    <li>
      <p>Nao also thought I should learn a text editor and introduced me to the one
he used: <a href="https://en.wikipedia.org/wiki/Vim_(text_editor)">Vim</a>. It is my primary text editor to this day (I’m writing
this post on it).</p>
    </li>
  </ul>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>Insight Data Science</h3></div>
    <div class="resume-location">Palo Alto, CA</div>
    <div class="resume-position">Data Scientist Fellow</div>
    <div class="resume-dates"><p>2015</p></div>
</div>

  <ul>
    <li>
      <p>I had decided not to go through the faculty lottery and was trying to figure
out my next steps. I was looking through my saved bookmarks one day when I
clicked on a blog written by my former student instructor at Berkeley,
<a href="https://twitter.com/berkeleyjess">Jessica Kirkpatrick</a>. The top post was <a href="https://berkeleyjess.blogspot.com/2014/07/career-profiles-astronomer-to-data.html"><em>Career Profiles:
Astronomer to Data Scientist</em></a>.<sup style="anchor-name:--fnref-jk_post" id="fnref:jk_post"><a href="#fn:jk_post" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Inside was a link to
Insight Data Science. There were just two days until applications to Insight
were due so I threw one together and was accepted.</p>
    </li>
    <li>
      <p>My advisor told me there was no way I would graduate by January, 2015 (he
was right). I emailed Insight and asked if I could delay starting until the
next session. They said yes. I have since learned their policy is not to do
that. I’m glad someone made an exception.</p>
    </li>
  </ul>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>Intuit</h3></div>
    <div class="resume-location">Mountain View, CA</div>
    <div class="resume-position">Staff Data Scientist</div>
    <div class="resume-dates"><p>2017–2020</p></div>
      <div class="resume-position">Senior Data Science Manager</div>
      <div class="resume-dates"><p>2018–2019</p></div>
</div>

  <ul>
    <li>My manager and director both left Intuit just after I joined, leaving our
team rudderless. It gave me an opportunity to step into a management role
first unofficially, and then officially several months later.</li>
  </ul>

  <div class="resume-header-grid">
    <div class="resume-company"><h3>Cash App</h3></div>
    <div class="resume-location">Remote</div>
    <div class="resume-position">Senior Staff (L7) Machine Learning Engineer, Modeler</div>
    <div class="resume-dates"><p>2023–Present</p></div>
      <div class="resume-position">Staff (L6) Machine Learning Engineer, Modeler</div>
      <div class="resume-dates"><p>2020–2023</p></div>
</div>

  <ul>
    <li>While <a href="/blog/interviewing-for-data-science-positions-in-2020/">looking for a job</a> after being laid off during COVID, I called
my friend and former coworker from Lab41, Patrick Callier, asking if he had
a lead on any positions at Square. He suggested I should work for Cash App
and introduced me to my current boss and skip-level (also both Insight
Alumni).</li>
  </ul>

</div>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:jk_post">

      <p>The text of <a href="https://berkeleyjess.blogspot.com/2014/07/career-profiles-astronomer-to-data.html"><em>Career Profiles: Astronomer to Data Scientist</em></a>:</p>

      <blockquote>
        <p><em>What, if any, additional training did you complete in order to meet the
qualifications?</em></p>

        <ol>
          <li>I participated in Scicoder where I learned about databases.</li>
          <li>I participated in a consulting internship where I learned about
working on interdisciplinary teams, tech/business applications of the
scientific method, and working with customers</li>
          <li>I was accepted to (but didn’t end up participating in) the <strong>Insight
Data Science Fellows Program</strong> where I would have learned more about the
transition from academia to tech, the tools used in data science /
analytics, and prepared for tech interviews. I got my job offer at
Yammer before this internship started so I participated as a
mentor/recruiter instead of a fellow.</li>
        </ol>
      </blockquote>
      <p><a href="#fnref:jk_post" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
      

      

      
      
        <summary type="html"><![CDATA[My career has been a great success so far and a lot of that success has been due to luck. In this post I catalog all the lucky breaks I can remember in the form of a resume.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/alt-resumes/landsknecht_playing_dice_by_theodor_alt_1913_sg468z.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/alt-resumes/landsknecht_playing_dice_by_theodor_alt_1913_sg468z.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Vim: I still love hjkl</title>
      <link href="https://alexgude.com/blog/i-still-love-hjkl/" rel="alternate" type="text/html" title="Vim: I still love hjkl" />
      <published>2022-05-16T00:00:00-07:00</published>
      <updated>2022-05-16T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/i_still_love_hjkl</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/i-still-love-hjkl/"><![CDATA[<p>I learned Vim in 2005 when I got my first job at Lawrence Berkeley Lab. I
needed to become comfortable working on the command line, which included
editing text files and scripts, so my mentor recommended the text editor he
used: <a href="https://www.vim.org">Vim</a>. It has been my primary editor ever since.<sup style="anchor-name:--fnref-neovim" id="fnref:neovim"><a href="#fn:neovim" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>Most Vim users love tweaking their editor configurations and perfecting their
editing habits. This is referred to as <a href="http://vimcasts.org/blog/2012/08/on-sharpening-the-saw/">sharpening their saws</a> by Drew
Neil. I am no exception of course—I made a <a href="/blog/vim-eldar/">custom color scheme,
<em>Eldar</em></a>, to make Vim that much more perfect for me—but there is one
area where I am an iconoclast:</p>

<p>I still love <code class="language-plaintext highlighter-rouge">hjkl</code>.</p>

<p>Let me explain:</p>

<h2 id="movement">Movement</h2>

<p>Moving around in Vim is a little different from other programs because there
are so many ways to do it. Most users start with the arrow keys, but they
quickly learn about using <code class="language-plaintext highlighter-rouge">hjkl</code> to move the cursor. Using <code class="language-plaintext highlighter-rouge">hjkl</code> might seem
unnatural but it starts the user down the path towards thinking in modes, and
besides it is nice to not have to move your hand off the home row.</p>

<p>Eventually, in the spirit of sharpening their saws, most users learn they can
use search, marks, and GOTOs to more efficiently “jump” to where they want to
go without mashing <code class="language-plaintext highlighter-rouge">jjjjjjjjj</code>. Some even argue that <code class="language-plaintext highlighter-rouge">hjkl</code> are irrelevant, as
done <a href="https://www.reddit.com/r/vim/comments/qh0zfz/comment/hia2xmy/?context=3">here on Reddit</a><sup style="anchor-name:--fnref-reddit_quote" id="fnref:reddit_quote"><a href="#fn:reddit_quote" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> or <a href="https://stackoverflow.com/a/26704213/1342354">here on Stack Overflow</a><sup style="anchor-name:--fnref-so_quote" id="fnref:so_quote"><a href="#fn:so_quote" class="footnote" rel="footnote" role="doc-noteref">3</a></sup>.</p>

<p>And they are right on one point: <code class="language-plaintext highlighter-rouge">hjkl</code> can’t compete on pure speed of
movement. But they are wrong when it comes to actually editing code.</p>

<h2 id="code-editing">Code Editing</h2>

<p>When I am editing code, I am not writing most of the time. I am reading. I am
thinking. I find <code class="language-plaintext highlighter-rouge">hjkl</code> amazingly effective for this.</p>

<p>Using <code class="language-plaintext highlighter-rouge">hjkl</code> lets me browse my code line by line, taking in the structure, the
layout, and considering what changes I am going to make. I feel more connected
with the logic of what the code describes than when I’m jumping quickly around
the file.</p>

<p>If you think about it, this is exactly the type of editing Vim was designed
for. Unlike most editors where you can just type and insert characters, Vim
forces you to decide that you are going to make an insertion and then switch
to insert mode to do it. Then you immediately leave insert mode. Vim expects
you to spend only part of your time writing, it expects you to spend a lot of
time reading, reorganizing, and navigating your code. And I find the lazy
scrolling of <code class="language-plaintext highlighter-rouge">hjkl</code> to fit perfectly with this philosophy.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:neovim">

      <p>At least until 2015, when I switched to <a href="https://neovim.io">Neovim</a>. Neovim is
a fork of Vim with the goal of creating a modern open source project that
is easy to contribute to, easy to maintain, and extend, all while adding
new features and making the editor an embeddable library. <a href="#fnref:neovim" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:reddit_quote">
      <p>Text of the Reddit comment:</p>

      <blockquote>
        <p>hjkl are irrelevant, it’s like micro movements, I want big fat movements
that put exactly where I need to be to do what I want.</p>

        <p>Each decision on how to navigate is informed by my intent for when I get
there. Snap decisions of course, I’m not sitting there working stuff out, it
just is automatic now.</p>

        <p>hjkl are just to help you do basic text editing. hjkl vs arrow keys is like
asking whether you prefer to crawl or drag yourself in a running race, just
learn to run.</p>
      </blockquote>
      <p><a href="#fnref:reddit_quote" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:so_quote">
      <p>Text of the Stack Overflow comment:</p>

      <blockquote>
        <p>The problem with the arrows is not that they are too far: the problem is
that they only allow you to move character-by-character and line-by-line.
And guess what? That is exactly what <code class="language-plaintext highlighter-rouge">hjkl</code> do. The only benefit of <code class="language-plaintext highlighter-rouge">hjkl</code>
over the arrows is that it saves that slight movement of the arm to and from
the arrows. Whether you think that benefit is worth the trouble is your
call. In my opinion, it isn’t.</p>

        <p><code class="language-plaintext highlighter-rouge">hjkl</code> are only <em>marginally better</em> than the arrows while Vim’s more
advanced motions, <code class="language-plaintext highlighter-rouge">bBeEwWfFtT,;/?^$</code> and so on, offer a <em>huge</em> advantage
over the arrows and <code class="language-plaintext highlighter-rouge">hjkl</code>.</p>

        <p>FWIW, I use the arrows for small movements, in normal and insert mode, and
the advanced motions above for larger motions.</p>

        <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>mouse-using sucker everyone laughs at:  (move)↓↓↓↓↓↓↓↓↓↓→→→→→(move)
hjkl-obsessed hipster:                        jjjjjjjjjjlllll efficient
vimmer:                             /fo&lt;CR&gt;
</code></pre></div>        </div>
      </blockquote>
      <p><a href="#fnref:so_quote" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[A lot of Vim users think you grow out of using hjkl for movement. After 17 years, I haven't. Here's why.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/hjkl/c_l_sholes_type_writing_machine_patent.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/hjkl/c_l_sholes_type_writing_machine_patent.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Comparing Zillow and Redfin Price Estimates in Time</title>
      <link href="https://alexgude.com/blog/online-realtor-estimate-timeseries/" rel="alternate" type="text/html" title="Comparing Zillow and Redfin Price Estimates in Time" />
      <published>2022-04-25T00:00:00-07:00</published>
      <updated>2022-04-25T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/online_realtor_estimate_timeseries</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/online-realtor-estimate-timeseries/"><![CDATA[<p>I <a href="/blog/online-realtor-estimate-comparison/">wrote about the pre- and post-sale estimates</a> of a house’s
price from four online brokers a little while ago. One limitation of that
experiment was that I only collected the data at a few points in time for the
house: one value before the listing, one value after the listing, and a final
value after the sale.</p>

<p>But what if the estimates change day-to-day? <a href="https://twitter.com/chrisemoody">Christopher Moody</a> posted
<a href="https://twitter.com/chrisemoody/status/1493686691378315264">this on Twitter</a>:</p>

<blockquote>
  <p>I suspect that their algos factor in impressions. I bet there’s a
fascinating time series in between pre- and post-listing price estimates</p>
</blockquote>

<p>I decided to look.</p>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Timeseries%20Plot.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Timeseries%20Plot.ipynb">rendered on Github</a>). The data can be found
<a href="/files/online-realtor-estimate-comparison//home_price_estimate_timeseries_data.csv">here</a>.</p>

<h2 id="data-collection">Data Collection</h2>

<p>I recorded the estimated sales price everyday from Zillow and Redfin. I
excluded Realtor.com and Xome because they do not update their estimate
frequently (they seem to only update when the status of the house changes from
pre-listing to listing to pending to sold). I occasionally missed a day of
data because I was collecting it manually.</p>

<p>Unfortunately the idea to collect the time series came after the house was
already pending, so I missed some of the most interesting changes. I have set
up a Python script to scrape listings for a few other houses and should have a
better time series for another article.</p>

<h2 id="plot">Plot</h2>

<p>Here a plot comparing the estimates for the sales price from Zillow and Redfin
in time:</p>

<p><a href="/files/online-realtor-estimate-comparison//home_price_estimate_timeseries_comparison.svg"><img src="/files/online-realtor-estimate-comparison//home_price_estimate_timeseries_comparison.svg" alt="A plot showing two time series, one is the estimated value of a house
according to Zillow, and the other is the estimate for the same house from
Redfin." /></a></p>

<p>The daily price estimates from Redfin are shown using red circles, the
estimates for Zillow are shown using blue triangles.</p>

<h3 id="comments">Comments</h3>

<p>There is <em>a lot</em> of movement in the estimates. Redfin and Zillow start off far
apart. Redfin reverts strongly to the list price once the house is official on
the market (I missed Zillow’s price on the day of listing). As the house goes
pending, both estimates increase drastically.</p>

<p>This suggests that some impression metric is used in both models, because
there is not much other information available that could cause such a large
increase. It will be interesting to see what a better time series will reveal.</p>

<p>Both estimates vary day-to-day even after the house is pending. You would
think all the information you need to determine the sale price would be
present once the offer is made, but there are facts you can observe after the
listing is pending—like how long the closing period is—that must have
predictive power. Redfin’s estimate remains pretty tight while Zillow’s
changes wildly, sometimes by 20% between days! The high variance really
reduces my confidence in Zillow’s model.</p>

<p>Finally, both estimates completely miss the actual sales price. The estimates
eventually jump to near the sales price, but it is not immediate; they take
about a week to adjust. Possibly a model update is triggered immediately for
newly listed properties—notice how the Redfin price adjusts the same day it
was listed—but not for sales? This makes some sense. A listing is an event
these business care about because they can profit off it, but a sale does not
provide that opportunity.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[How do Zillow and Redfin's home price estimates change in time? I use a recently sold house in my neighborhood to find out. Come check it out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_houses_in_san_francisco.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_houses_in_san_francisco.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Gatekeeper Teams of the Overwatch League</title>
      <link href="https://alexgude.com/blog/overwatch-league-gatekeeper-teams/" rel="alternate" type="text/html" title="Gatekeeper Teams of the Overwatch League" />
      <published>2022-03-07T00:00:00-08:00</published>
      <updated>2022-03-07T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/overwatch_league_gatekeeper_teams</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/overwatch-league-gatekeeper-teams/"><![CDATA[<p>Watching the <a href="https://en.wikipedia.org/wiki/Overwatch_League">Overwatch league</a> is probably my nerdiest hobby (well after
doing data analysis on the weekend so I can write these posts of course).</p>

<p>Overwatch fans love debating which team each year is the “gatekeeper”. The
gatekeeper is the team that beats lower ranked teams, often in a lopsided
fashion, but can’t themselves move out of the <a href="https://en.wiktionary.org/wiki/midtable">midtable</a>. In this way
they gatekeep the standings: top teams can beat them, bottom teams can’t, and
so they define the dividing line between the two groups.</p>

<p>Determining which team is the gatekeeper of each season seems like a good
question to answer with data. I can look for a team that is good at beating
low-level competitors, but fails against the top teams.</p>

<h2 id="gatekeeper-score">Gatekeeper Score</h2>

<p>To find the gatekeepers of each season, I need to define some metric to
measure them by. I will use what I call the <strong>gatekeeper score</strong>.</p>

<p>The gatekeeper score is the win percentage against teams that finished
lower in the regular season standings minus the win percentage against teams
that finished higher in the standings. A perfect gatekeeper team, one that
beats all lower ranked teams but loses to all the teams above them, would have
a gatekeeper score of 100 (100% win rate against lower teams minus 0% win rate
against high teams).</p>

<p>This score captures most of what I want, but there are a few problems:</p>

<h3 id="undefined-scores">Undefined Scores</h3>

<p>The best and worst team each season have an undefined score because there are
no teams better or worse than them. This is not a big problem because to be a
gatekeeper you must gatekeep someone, and you can’t do that at the top or
bottom of the standings.</p>

<h3 id="high-volatility-near-the-top-and-bottom">High Volatility Near the Top and Bottom</h3>

<p>Teams near the top or the bottom of the standings have a small number of
matches used to compute one component of their score. This tends to increase
the volatility of their scores relative to midtable teams that have a lot of
matches counted on either side.</p>

<p>This is a bigger problem because it means that teams near the top and bottom
are more likely to have a high gatekeeper score just because they won or lost
a single match, whereas teams in the middle need to win or lose many matches
to change their score.</p>

<p>I could adjust the score by the number of games, or compute an estimate of the
variance, but for now I will just call out this issue as I run into it.</p>

<h2 id="the-data">The Data</h2>

<p>I used two sources of data for this comparison. The first is <a href="https://overwatchleague.com/en-us/statslab">a record of the
outcome of every map played</a> from the Overwatch League’s Stat Lab.
I transform this data to get match-level<sup style="anchor-name:--fnref-match" id="fnref:match"><a href="#fn:match" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> win-loss records for each
team by opponent. You can find that data <a href="/files/overwatch_league//match_level_data.json">here</a> and the
notebook to parse it is <a href="/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">rendered on
Github</a>).</p>

<p>The second data source is the regular season standings of all the teams. I
scrape this data from <a href="https://liquipedia.net/overwatch/Overwatch_League">Liquipedia</a>. I used the regular season
standings because the final standings are based on a handful of playoff games
while the season standings incorporate many more matches and so provide a more
accurate estimate of a team’s performance. The parsed standings data is
<a href="/files/overwatch_league//owl_standings.json">here</a>. The code to generate the data frame is
<a href="/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">rendered on Github</a>).</p>

<p>To compute the gatekeeper score I used both regular season and tournament
games, primarily because the dataset does not separate them and I don’t want
to go label them by hand.</p>

<p>The final, combined data frame, can be found <a href="/files/overwatch_league//combined_standings_data.json">here</a>. The
notebook to read it is <a href="/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/overwatch_league//OWL%20Gatekeeper%20Teams.ipynb">rendered on
Github</a>).</p>

<h2 id="the-gatekeepers">The Gatekeepers</h2>

<p>Here are the gatekeeper scores for each season and region. The tables are
ordered according to the regular season ranking. I have highlighted teams I
think could rightfully be called gatekeepers.</p>

<h3 id="2018-season">2018 Season</h3>

<table>
  <thead>
    <tr>
      <th>Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>New York Excelsior</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td>Los Angeles Valiant</td>
      <td style="text-align: right">28</td>
    </tr>
    <tr>
      <td>Boston Uprising</td>
      <td style="text-align: right">8</td>
    </tr>
    <tr>
      <td>Los Angeles Gladiators</td>
      <td style="text-align: right">56</td>
    </tr>
    <tr>
      <td>London Spitfire</td>
      <td style="text-align: right">23</td>
    </tr>
    <tr>
      <td>Philadelphia Fusion</td>
      <td style="text-align: right">33</td>
    </tr>
    <tr>
      <td>Houston Outlaws</td>
      <td style="text-align: right">28</td>
    </tr>
    <tr>
      <td><strong>Seoul Dynasty</strong></td>
      <td style="text-align: right"><strong>61</strong></td>
    </tr>
    <tr>
      <td>San Francisco Shock</td>
      <td style="text-align: right">54</td>
    </tr>
    <tr>
      <td>Dallas Fuel</td>
      <td style="text-align: right">51</td>
    </tr>
    <tr>
      <td><strong>Florida Mayhem</strong></td>
      <td style="text-align: right"><strong>89</strong></td>
    </tr>
    <tr>
      <td>Shanghai Dragons</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>The <a href="https://en.wikipedia.org/wiki/2018_Florida_Mayhem_season">2018 Florida Mayhem</a> were a bad team with a possibly even
worse uniform. They went 7-33 in the inaugural season and were saved from the
bottom of the rankings by the <strong>team with the</strong> <a href="https://www.espn.com/esports/story/_/id/25535277/espn-esports-awards-2018-why-shanghai-dragons-0-40-record-espn-biggest-disappointment-year"><strong>worst losing
streak</strong></a> <strong>in professional sports history:</strong> the <a href="https://en.wikipedia.org/wiki/2018_Shanghai_Dragons_season">0-40 Shanghai
Dragons</a>.</p>

<p>The Mayhem went 3-0 against the Dragons for a 100% win rate against worse
teams, and 4-33 against better teams giving them an impressively bad 11% win
rate. Combining those rates yields a gatekeeper score of 89!</p>

<p>But are they the gatekeeper? They run into the <a href="#high-volatility-near-the-top-and-bottom">volatility problem I described
above</a>: 3 wins against the Dragons got them 100 points and 4 wins
against better teams lost them only 11.</p>

<p>I think the <a href="https://en.wikipedia.org/wiki/2018_Seoul_Dynasty_season">2018 Seoul Dynasty</a> are a better candidate for the
gatekeeper team. They went 22-18, solidly middle-of-the-pack, and they have a
93% win rate against lower ranked teams with only a 32% win rate against
higher rated teams. They are also the first team with a positive win rate and
map differential as you work your way up from the bottom.</p>

<h3 id="2019-season">2019 Season</h3>

<p>The 2019 season added eight teams to the league and was dominated by the
highly-technical <a href="https://thegamehaus.com/overwatch/a-comprehensive-history-of-overwatch-metas-part-15-goats/2020/02/06/">GOATS meta</a> in which teams played three tanks and
three supports. During this meta, teams were forced to rigorously track their
opponents ability usage and perfectly time their own in order to win fights.
The Vancouver Titans, the San Francisco Shock, and to a lesser extent the New
York Excelsior mastered this style of play and dominated the league while many
other teams failed to achieve the high-level of coordination required and sunk
down in the rankings.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Vancouver Titans</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td style="text-align: left">San Francisco Shock</td>
      <td style="text-align: right">24</td>
    </tr>
    <tr>
      <td style="text-align: left">New York Excelsior</td>
      <td style="text-align: right">56</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Hangzhou Spark</strong></td>
      <td style="text-align: right"><strong>75</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Los Angeles Gladiators</td>
      <td style="text-align: right">44</td>
    </tr>
    <tr>
      <td style="text-align: left">Atlanta Reign</td>
      <td style="text-align: right">12</td>
    </tr>
    <tr>
      <td style="text-align: left">London Spitfire</td>
      <td style="text-align: right">39</td>
    </tr>
    <tr>
      <td style="text-align: left">Seoul Dynasty</td>
      <td style="text-align: right">47</td>
    </tr>
    <tr>
      <td style="text-align: left">Guangzhou Charge</td>
      <td style="text-align: right">46</td>
    </tr>
    <tr>
      <td style="text-align: left">Philadelphia Fusion</td>
      <td style="text-align: right">28</td>
    </tr>
    <tr>
      <td style="text-align: left">Shanghai Dragons</td>
      <td style="text-align: right">13</td>
    </tr>
    <tr>
      <td style="text-align: left">Chengdu Hunters</td>
      <td style="text-align: right">38</td>
    </tr>
    <tr>
      <td style="text-align: left">Los Angeles Valiant</td>
      <td style="text-align: right">26</td>
    </tr>
    <tr>
      <td style="text-align: left">Paris Eternal</td>
      <td style="text-align: right">25</td>
    </tr>
    <tr>
      <td style="text-align: left">Dallas Fuel</td>
      <td style="text-align: right">55</td>
    </tr>
    <tr>
      <td style="text-align: left">Houston Outlaws</td>
      <td style="text-align: right">26</td>
    </tr>
    <tr>
      <td style="text-align: left">Toronto Defiant</td>
      <td style="text-align: right">7</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Washington Justice</strong></td>
      <td style="text-align: right"><strong>83</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Boston Uprising</td>
      <td style="text-align: right">24</td>
    </tr>
    <tr>
      <td style="text-align: left">Florida Mayhem</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>Once again a bottom-ranked team has the highest score and for the same reason:
<a href="#high-volatility-near-the-top-and-bottom">volatility</a>. The <a href="https://en.wikipedia.org/wiki/2019_Washington_Justice_season">2019 Washington Justice</a> went 4-0
against Boston and Florida and 4-20 against better teams.</p>

<p>More interesting are the <a href="https://en.wikipedia.org/wiki/2019_Hangzhou_Spark_season">2019 Hangzhou Spark</a>, who placed fourth.
They lost every game against the top three teams, going 0-7, including 2 lost
playoff games against the eventual champions the San Francisco Shock. But the
Spark made up for it with a 75% win rate against worse teams.</p>

<p>What feels right about the highly-ranked Spark being the gatekeepers is it
follows the storyline of the 2019 season: that the Vancouver Titans and the
San Francisco Shock were a tier above everyone else in the GOATS meta, that
only a few other top teams were able to execute the GOATS composition with
enough coordination (among them New York), and that everyone else floundered
trying to play a meta they did not have the skill to.</p>

<h3 id="2020-season">2020 Season</h3>

<p>The 2020 Overwatch League season was split into two regions with very few
games played between them due to the <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">2020 COVID pandemic</a>. For that
reason I have split the comparison in two and look at the North American and
Asian regions separately.</p>

<h4 id="north-america">North America</h4>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Philadelphia Fusion</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td style="text-align: left">San Francisco Shock</td>
      <td style="text-align: right">14</td>
    </tr>
    <tr>
      <td style="text-align: left">Paris Eternal</td>
      <td style="text-align: right">31</td>
    </tr>
    <tr>
      <td style="text-align: left">Florida Mayhem</td>
      <td style="text-align: right">41</td>
    </tr>
    <tr>
      <td style="text-align: left">Los Angeles Valiant</td>
      <td style="text-align: right">15</td>
    </tr>
    <tr>
      <td style="text-align: left">Los Angeles Gladiators</td>
      <td style="text-align: right">50</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Atlanta Reign</strong></td>
      <td style="text-align: right"><strong>67</strong></td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Dallas Fuel</strong></td>
      <td style="text-align: right"><strong>77</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Toronto Defiant</td>
      <td style="text-align: right">62</td>
    </tr>
    <tr>
      <td style="text-align: left">Houston Outlaws</td>
      <td style="text-align: right">6</td>
    </tr>
    <tr>
      <td style="text-align: left">Vancouver Titans</td>
      <td style="text-align: right">39</td>
    </tr>
    <tr>
      <td style="text-align: left">Washington Justice</td>
      <td style="text-align: right">69</td>
    </tr>
    <tr>
      <td style="text-align: left">Boston Uprising</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>The <a href="https://en.wikipedia.org/wiki/2020_Atlanta_Reign_season">2020 Atlanta Reign</a> were the first time I remember a team
specifically being referred to as a gatekeeper. They had flashy, lopsided
games against teams below them in the standings, but constantly failed to
advance by beating teams ahead of them. You can see it in their stats: the
Reign have an 85% win rate against lower ranked teams—the third highest in
the league behind the eventual champions the Shock and Washington who only had
team below them—but just an 18% win rate against higher ranked teams.</p>

<p>However, I think the <a href="https://en.wikipedia.org/wiki/2020_Dallas_Fuel_season">2020 Dallas Fuel</a> are the true gatekeepers of
the league. They had almost as high a win rate against worse teams at 83%, but
a much lower win rate against better teams at 7%—the lowest in the league
that year!</p>

<h4 id="asia">Asia</h4>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Shanghai Dragons</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td style="text-align: left">Guangzhou Charge</td>
      <td style="text-align: right">40</td>
    </tr>
    <tr>
      <td style="text-align: left">New York Excelsior</td>
      <td style="text-align: right">38</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Hangzhou Spark</strong></td>
      <td style="text-align: right"><strong>50</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Seoul Dynasty</td>
      <td style="text-align: right">32</td>
    </tr>
    <tr>
      <td style="text-align: left">Chengdu Hunters</td>
      <td style="text-align: right">23</td>
    </tr>
    <tr>
      <td style="text-align: left">London Spitfire</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>There were two facts in the 2020 Asian region: the Shanghai Dragons were the
best team in the region by a mile, and the London<sup style="anchor-name:--fnref-london" id="fnref:london"><a href="#fn:london" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> Spitfire were the
worst by a wide margin.</p>

<p>Amongst the remaining teams, the <a href="https://en.wikipedia.org/wiki/2020_Hangzhou_Spark_season">2020 Hangzhou Spark</a> look like a
good candidate for the gatekeepers. They beat lower ranked teams
67% of the time, but only managing 17% against better teams, and they finished
exactly in the middle of the pack in the standings.</p>

<h2 id="2021-season">2021 Season</h2>

<p>The 2021 season was again split into two regions due to COVID, but there was a
lot more cross-play between the regions as each of the four tournaments and
the playoffs included teams from both regions. Still, I only look at the
record against teams within the region as that is how regular season standings
were determined.</p>

<h3 id="north-america-1">North America</h3>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Dallas Fuel</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Los Angeles Gladiators</strong></td>
      <td style="text-align: right"><strong>70</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Atlanta Reign</td>
      <td style="text-align: right">32</td>
    </tr>
    <tr>
      <td style="text-align: left">San Francisco Shock</td>
      <td style="text-align: right">43</td>
    </tr>
    <tr>
      <td style="text-align: left">Houston Outlaws</td>
      <td style="text-align: right">47</td>
    </tr>
    <tr>
      <td style="text-align: left">Washington Justice</td>
      <td style="text-align: right">39</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Toronto Defiant</strong></td>
      <td style="text-align: right"><strong>62</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Paris Eternal</td>
      <td style="text-align: right">46</td>
    </tr>
    <tr>
      <td style="text-align: left">Boston Uprising</td>
      <td style="text-align: right">31</td>
    </tr>
    <tr>
      <td style="text-align: left">Florida Mayhem</td>
      <td style="text-align: right">69</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>London Spitfire</strong></td>
      <td style="text-align: right"><strong>100</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Vancouver Titans</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>The <a href="https://en.wikipedia.org/wiki/2021_London_Spitfire_season">2021 London Spitfire</a> were a disappointing team. The
majority of the players had been called up from the Spitfire’s academy team
the <a href="https://en.wikipedia.org/wiki/British_Hurricane">British Hurricane</a> who had gone 12-0 and won the 2020
Overwatch Contenders season (Overwatch’s minor league). But they floundered in
the Overwatch league, barely scrapping together a 1-15 season. They got their
only win in the infamous <a href="https://overwatchleague.com/en-us/match/37333">Bread Bowl</a> against the 1-15
<a href="https://en.wikipedia.org/wiki/2021_Vancouver_Titans_season">Vancouver Titans</a>. That gave the Spitfire a 1-0 record against
lower ranked teams and a 0-15 record against higher ranked teams, resulting in
a perfect 100 point gatekeeper score.</p>

<p>Of course a one-win team isn’t a gatekeeper, despite what the score says.
Likewise it’s hard to call the <a href="https://en.wikipedia.org/wiki/2021_Los_Angeles_Gladiators_season">2021 Los Angeles Gladiators</a>
gatekeeper because their 0% win rate against higher ranked teams (just the
<a href="https://en.wikipedia.org/wiki/2021_Dallas_Fuel_season">Dallas Fuel</a>) represents only two games.</p>

<p>For that reason, I think the <a href="https://en.wikipedia.org/wiki/2021_Toronto_Defiant_season">2021 Toronto Defiant</a> are the
gatekeepers of the North American region. They had a typical middle of the
pack season: finishing 7 out of 12 in the West, a 9-7 match record, and a
perfectly balanced map record with 32 wins and 32 losses. To round it out, they
had an 82% win rate against worse teams and just a 20% win rate against better
teams.</p>

<h3 id="asia-1">Asia</h3>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Team</th>
      <th style="text-align: right">Gatekeeper Score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Shanghai Dragons</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td style="text-align: left">Chengdu Hunters</td>
      <td style="text-align: right">53</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Seoul Dynasty</strong></td>
      <td style="text-align: right"><strong>62</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Philadelphia Fusion</td>
      <td style="text-align: right">40</td>
    </tr>
    <tr>
      <td style="text-align: left">Hangzhou Spark</td>
      <td style="text-align: right">42</td>
    </tr>
    <tr>
      <td style="text-align: left">New York Excelsior</td>
      <td style="text-align: right">27</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Guangzhou Charge</strong></td>
      <td style="text-align: right"><strong>79</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Los Angeles Valiant</td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>The <a href="https://en.wikipedia.org/wiki/2021_Guangzhou_Charge_season">2021 Guangzhou Charge</a> have the highest gatekeeper score,
but it is, once again, <a href="#high-volatility-near-the-top-and-bottom">due to only having a single team below
them</a>: the winless <a href="https://en.wikipedia.org/wiki/2021_Los_Angeles_Valiant_season">2021 Los Angeles Valiant</a>. So
I do not think they’re a good choice for gatekeepers.</p>

<p>Instead, I would choose the <a href="https://en.wikipedia.org/wiki/2021_Seoul_Dynasty_season">2021 Seoul Dynasty</a>, who had a
similarly bad win rate against better teams (22% for Dynasty, 21% for
Charge), but who earn their 85% win rate against lower ranked teams honestly
by beating teams that have actually won games, like the 10 and 10 <a href="https://en.wikipedia.org/wiki/2021_Philadelphia_Fusion_season">2021
Philadelphia Fusion</a>.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:match">

      <p>A match consists of multiple maps that are played sequentially. The
first team to win a specific number of maps wins the match. The
number of map wins needed is often 3, but occasionally 4 or more for
tournaments. Some seasons all maps were played out even if one team
had already clinched the match. <a href="#fnref:match" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:london">

      <p>I know London is not in Asia, and neither is New York, but during
the pandemic these teams decided to move to Korea for safety. <a href="#fnref:london" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="fun-and-games" />
        
      

      

      
      
        <summary type="html"><![CDATA[In the Overwatch League, a "gatekeeper" team is one that beats all the lower ranked teams but can't beat higher ranked teams. I use match data to determine which teams are each season's gatekeeper.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/overwatch_league/gate_of_damascus_jerusalem_april_14_1939_by_louis_haghe_and_david_roberts.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/overwatch_league/gate_of_damascus_jerusalem_april_14_1939_by_louis_haghe_and_david_roberts.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Comparing Pre- and Post-sale Estimates of the Price of a House</title>
      <link href="https://alexgude.com/blog/online-realtor-estimate-comparison/" rel="alternate" type="text/html" title="Comparing Pre- and Post-sale Estimates of the Price of a House" />
      <published>2022-02-07T00:00:00-08:00</published>
      <updated>2022-02-07T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/online_realtor_estimate_comparison</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/online-realtor-estimate-comparison/"><![CDATA[<p>When I bought my house several years ago I was unsure how much money it would
take to actually buy it. There was the list-price, of course, but I didn’t
expect that to be the price the seller would accept. Zillow and Redfin
estimated two different prices, both slightly higher than the list price, and
my realtor suggested a fourth based on comparable homes. In the end we made up a
number combining all four prices, nudged it up a bit as a hedge, and had our
offer accepted. It left me wondering if there was a better way to predict
prices, and if Zillow and Redfin had been right with their predicted prices.</p>

<p>Recently a house in my neighborhood put up a “For Sale” sign which prompted me
to look online for the listing. I couldn’t find one. None of the online real
estate brokers had picked it up yet. I realized I had a chance to compare
their current price estimates with the actual listing and sales price.</p>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Plot.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/online-realtor-estimate-comparison//House%20Price%20Estimate%20Plot.ipynb">rendered on Github</a>). The data can be found
<a href="/files/online-realtor-estimate-comparison//house_price_estimate_date.csv">here</a>.</p>

<h2 id="data-collection">Data Collection</h2>

<p>I recorded the pre-listing estimates for four brokers: Redfin, Realtor.com,
Zillow, and Xome. Xome provides both an estimate and a high and low range. All
the others only present a single estimate. I also collected the same estimates
after the listing was picked up and again after the sale was complete. The
data is summarized in the table below:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Company</th>
      <th style="text-align: right">Pre-listing</th>
      <th style="text-align: right">Post-listing</th>
      <th style="text-align: right">Post-sale</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><strong>Zillow</strong></td>
      <td style="text-align: right">$938K</td>
      <td style="text-align: right">$941K</td>
      <td style="text-align: right">$1077K</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Realtor.com</strong></td>
      <td style="text-align: right">$977K</td>
      <td style="text-align: right">Not Recorded</td>
      <td style="text-align: right">$1105K</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Xome</strong></td>
      <td style="text-align: right">$1040K<span class="supsub"><sup>+90K</sup><sub>-91K</sub></span></td>
      <td style="text-align: right">Unchanged</td>
      <td style="text-align: right">$1074K<span class="supsub"><sup>+91K</sup><sub>-113K</sub></span></td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Redfin</strong></td>
      <td style="text-align: right">$1144K</td>
      <td style="text-align: right">$963K</td>
      <td style="text-align: right">$1090K</td>
    </tr>
  </tbody>
</table>

<p>I could not find Realtor.com’s estimate after the listing went up, so it is
not included. Xome did not change their estimate after the listing was posted.</p>

<p>The house was listed at $948K and sold for $1070K.</p>

<h2 id="plot">Plot</h2>

<p>Here a plot comparing the three estimates from each company to the list and
sales price:</p>

<p><a href="/files/online-realtor-estimate-comparison//home_price_estimate_comparison.svg"><img src="/files/online-realtor-estimate-comparison//home_price_estimate_comparison.svg" alt="A comparison of four different real estate brokers estimate for the sales
price of a single house in my neighborhood before and after it was listed and
sold." /></a></p>

<p>The three price estimates are:</p>

<ul>
  <li>
    <p><strong>Pre-listing</strong>: Before the listing was picked up by the brokers,
represented with a circle: ●</p>
  </li>
  <li>
    <p><strong>Post-listing</strong>: After the listing was posted but before the sale,
represented with a triangle: ▲</p>
  </li>
  <li>
    <p><strong>Post-sale</strong>: After the sale was made public, represented with a square:
■</p>
  </li>
</ul>

<p>I have slightly offset the date points for each company—with the circle on
left, the triangle in the middle, and the square on the right—to give a
quick indication of how the price trended in time if you read from left to
right. The Xome estimates include error bars for their high and low estimates.
The listing and sale price are shown as lines. I have colored each company’s
estimates to match their brand color. The companies are sorted from lowest to
highest pre-listing estimate.</p>

<h3 id="comments">Comments</h3>

<p>Zillow and Redfin <em>strongly</em> disagree about what the value of the house is
initially, with a difference between their estimates of $200K. Both estimates
poorly predict the final sales price, with Zillow underestimating by $132K and
Redfin overestimating by $74K.</p>

<p>Both estimates revert towards the list price when it is posted, with Redfin’s
slightly higher than the listing and Zillow’s slightly lower. This makes some
sense, as the list price contains new information about the current market
conditions and more importantly about the condition of the house and property
relative to its neighbors. However, in this case the list price was obviously
too low and likely intended to entice buyers.</p>

<p>Xome’s pre-list estimate is the closest to the sale value, missing by just
about 3%, although they gave themselves a lot of room with their uncertainty.
Here are the four pre-listing estimates ranked from lowest to highest absolute
error:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Company</th>
      <th style="text-align: right">Absolute Error</th>
      <th style="text-align: right">Percent Error</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><strong>Xome</strong></td>
      <td style="text-align: right">$30K</td>
      <td style="text-align: right">2.8%</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Redfin</strong></td>
      <td style="text-align: right">$74K</td>
      <td style="text-align: right">6.9%</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Realtor.com</strong></td>
      <td style="text-align: right">$93K</td>
      <td style="text-align: right">8.7%</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Zillow</strong></td>
      <td style="text-align: right">$132K</td>
      <td style="text-align: right">12.3%</td>
    </tr>
  </tbody>
</table>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[Can Zillow and Redfin predict prices accurately? I look at a house sold in my neighborhood and compare the sale price to the price predicted by Zillow and Redfin before they knew it was for sale.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_house_plans.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/online-realtor-estimate-comparison/1935_house_plans.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Comparison of My Three Sons’ Language Development</title>
      <link href="https://alexgude.com/blog/all-my-sons-language-development-comparison/" rel="alternate" type="text/html" title="Comparison of My Three Sons’ Language Development" />
      <published>2022-01-03T00:00:00-08:00</published>
      <updated>2022-01-03T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/all_my_sons_language_development_comparison</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/all-my-sons-language-development-comparison/"><![CDATA[<p>I tracked the language development of all three of my sons. I wrote <a href="/blog/my-sons-words/">a post
focusing on Theo’s language development</a>, <a href="/blog/my-second-sons-words/">another post focusing
on Cory’s language development</a>, and <a href="/blog/my-third-sons-words/">a final one focusing on Ash’s
language development</a>. In this post I’ll compare them all.</p>

<h2 id="the-data">The Data</h2>

<p>The data was collected by my wife and I using a Google form on our phones.
When we heard a new word we would log it. I then normalized the words
(sometimes we would write down grandma and sometimes grandmother for the same
word, for example) by hand and took the first occurrence of each word in each
language as the date when they learned it.</p>

<p>I discuss data collection in more depth in <a href="/blog/my-sons-words/#the-data">Theo’s</a>,
<a href="/blog/my-second-sons-words/#the-data">Cory’s</a>, and <a href="/blog/my-third-sons-words/#the-data">Ash’s</a> data sections. You can
find the Jupyter notebook used to perform this analysis <a href="/files/all-my-sons-words-comparison//Theo%20vs%20Cory%20vs%20Ash%20words.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/all-my-sons-words-comparison//Theo%20vs%20Cory%20vs%20Ash%20words.ipynb">rendered on Github</a>). The data can be found <a href="/files/my-sons-words/theo_words.csv">here</a>,
<a href="/files/my-second-sons-words/cory_words.csv">here</a>, and <a href="/files/my-third-sons-words/ash_words.csv">here</a>.</p>

<h2 id="development">Development</h2>

<p>Below I have plotted the number of words each of my sons knew as a function of
their age. Theo, our first son, is represented by dotted lines; Cory, our
second son, is represented by dashed lines; Ash, our third son, is represented
with solid lines.</p>

<p><a href="/files/all-my-sons-words-comparison//child0_vs_child1_vs_child2_total_words_linear.svg"><img src="/files/all-my-sons-words-comparison//child0_vs_child1_vs_child2_total_words_linear.svg" alt="A plot showing the number of words my sons could speak as a function of
their age." /></a></p>

<p>Cory learned all of the languages much faster than his brothers. Theo was the
slowest to learn English and Chinese, the primary languages of our household,
but slightly faster than Ash on Spanish, Sign, and Animal Sounds.
Interestingly, Ash’s Chinese started strong but has slowed; had we kept
recording I suspect Theo would have caught up and passed him at 26 months of
age.</p>

<p>I discuss the slow Chinese acquisition in the <a href="/blog/my-third-sons-words/#development">development
section of Ash’s post</a>, but briefly I think it is related to me
working from home during the COVID pandemic. This meant that:</p>

<ul>
  <li>
    <p>We used a lot more English at home.</p>
  </li>
  <li>
    <p>I was able to better identify when he had learned new English words but not
new Chinese words, which led to a selection effect in what was recorded.</p>
  </li>
</ul>

<p>With one additional month of living with Ash (not reflected in the chart and
data) his Chinese has really taken off recently, using whole sentences instead
of just a few words, which leads me to suspect that the slowdown in the data
is mostly due to the selection effect mentioned above.</p>

<p>Also interestingly my wife tells me that Cory’s current Chinese is worse than
Theo’s was at the same age. Perhaps Cory’s development slowed or there is more
to proficiency than just number of words known.</p>

<p>The shape of the English curves are all very similar, just displaced in time.
I would guess this has to do with the fact that they are always surrounded by
English speakers to learn from whereas the other languages require special
interactions with their family to learn.</p>

<p>Ash’s Spanish has taken off much slower than our other two sons, although in
comparison he is not too far behind Theo. Theo’s long drought of Spanish words
is due to an injury my mom sustained that prevented her and my father from
watching Theo and so reduced his contact with Spanish. Ash doesn’t have the
same excuse, but it does show it’s not too late for him to pick it up.</p>

<p>Ash never got much into sign language. He learned pointing quickly and got by
just doing that until he could speak. We also didn’t emphasize it as much
because we were very busy with our other two boys. I suspect this is why he
didn’t know as many animal sounds as well: I used to sit with Theo and Cory
and point at animals in books and teach them but I did not do this with Ash.</p>

<p>As a final note: our kids’ doctors were worried about both Theo and Ash’s
language development being too slow. Some of that worry transferred to us as
parents. Keeping track of their words like this was reassuring—we could see
that Ash was on pace compared to Theo, and we knew Theo turned out fine!</p>

<h2 id="other-writings-on-language-development">Other Writings on Language Development</h2>

<p>If you enjoyed this article, here are all the other articles I wrote about
<a href="/topics/childhood-language/">language development</a>!</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <img src="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" alt="A woodcut by J. W. Orr showing a father giving his son a picture book in a richly appointed study.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <strong>My Third Son’s Language Development</strong>
    </a>
<br />
We tracked my third son's language development word by word. Here, in plots, is how he learned to speak. Take a look!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" alt="Black and white photo of a young boy at a school desk.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <strong>Comparison of My Two Sons’ Language Development</strong>
    </a>
<br />
Being a nerd dad, I recorded all the words my first two sons spoke as they learned them. Now, I compare their language development rate!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <img src="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a man using a blackboard to teach young children punctuation.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <strong>My Second Son’s Language Development</strong>
    </a>
<br />
My second son is a little over two years old. We tracked every word he's spoken to watch his language development, and now you can observe it too!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <img src="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a woman using a blackboard to teach young children how to pronounce words.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <strong>My Son’s Language Development</strong>
    </a>
<br />
My son is a little over two and unfortunately he has two huge nerds for parents. We tracked every word he's spoken to watch his language development, and now you can join us!  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="childhood-language" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[I recorded the words my sons spoke as they learned our various languages and now I compare how each developed! Read on to find out how each son learned.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Third Son’s Language Development</title>
      <link href="https://alexgude.com/blog/my-third-sons-words/" rel="alternate" type="text/html" title="My Third Son’s Language Development" />
      <published>2021-12-20T00:00:00-08:00</published>
      <updated>2021-12-20T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/my_third_sons_words</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-third-sons-words/"><![CDATA[<p>My Third son, Ash, was born in the fall of 2019. Like my <a href="/blog/my-sons-words/">first</a>
and <a href="/blog/my-second-sons-words/">second</a> sons, we tracked Ash’s language development to monitor
how quickly he was picking up the different languages we speak.</p>

<h2 id="the-data">The Data</h2>

<p>The data was collected in the <a href="/blog/my-sons-words/#the-data">same manner as last two times</a>.
One difference was that Ash learned to speak during the <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">COVID-19
pandemic</a>, when we were much more isolated from family and the rest
of the world, but also I was working from home so it was easier to enter my
own data rather than have my wife text me when Ash said a new word.</p>

<p>The main difficulties in data collection were the same though:</p>

<ul>
  <li>
    <p>Deciding when Ash had associated a sound with a concept as opposed to just
babbling.</p>
  </li>
  <li>
    <p>Deciding if Ash “knew” a word or was just repeating a sound he had just
heard.</p>
  </li>
</ul>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/my-third-sons-words//Ash's%20first%20words.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/my-third-sons-words//Ash's%20first%20words.ipynb">rendered on Github</a>). The data can be found
<a href="/files/my-third-sons-words//ash_words.csv">here</a>.</p>

<h2 id="development">Development</h2>

<p>Ash’s first word was “milk” in Cantonese, spoken at 13 months old. Milk was
something he loved back then and the word is still in frequent use. He often
demands “milk milk milk” while shoving his empty cup towards us. His second
word was “mom” in Cantonese and his third word was “dad” in English. Unlike my
other children, Ash didn’t start using the Cantonese word for “dad” until very
late. He preferred his own <a href="https://en.wikipedia.org/wiki/Chinglish">pidgin</a> where he would simply use the
English word “dad” in an otherwise Cantonese phrase.</p>

<p>Ash learned a few words of <a href="https://en.wikipedia.org/wiki/Baby_sign_language">baby sign language</a>, but unlike his
other brothers was not too interested in learning much.</p>

<p>Ash’s language development is plotted below, showing the number of words he
could speak in each “language” as a function of how old he was.</p>

<p><a href="/files/my-third-sons-words//child2_total_words_linear.svg"><img src="/files/my-third-sons-words//child2_total_words_linear.svg" alt="A plot showing the number of words my third son could speak as a function
of age." /></a></p>

<p>Ash picked up English and Cantonese at roughly the same rate, with Cantonese
leading in the number of words spoken until he was 23 months old when he
started speaking more English. I suspect the reason his English learning
outpaced his Cantonese is because I was working from home during the pandemic.
This had two effects:</p>

<ul>
  <li>
    <p>Ash got a lot more exposure to my wife and other sons speaking to me in
English, and more time listening to me talk to him.</p>
  </li>
  <li>
    <p>I was able to record new English words he was speaking, but since my
Cantonese is bad (really bad) I could not do the same for it and so relied
on my wife. This made it more likely that I would write down a new English
word while missing more Cantonese words.</p>
  </li>
</ul>

<p>Ash’s language development really took off at 17 or 18 months of age and went
nearly vertical at 23 months. During his 23rd month he doubled the number of
English words he knew to about 100 and overtook the number of Cantonese words,
which climbed more slowly.</p>

<p>Ash’s Spanish development has been slow. He has had plenty of time with my
parents (who both speak Spanish) but not much alone time with them. He
normally visits with his older brothers who speak mostly English to my parents
now, which has reduced Ash’s exposure to Spanish. Ash also did not love
Spanish cartoons, which his older brothers did at his age.</p>

<h2 id="the-words">The Words</h2>

<p>I plotted a selection of some of Ash’s first words in each language below.
Notice that I have switched to a log plot for the <em>y</em>-axis to better show the
beginnings of each language.</p>

<p><a href="/files/my-third-sons-words//child2_first_words.svg"><img src="/files/my-third-sons-words//child2_first_words.svg" alt="A plot showing the third words my son could speak as a function of
age." /></a></p>

<p>Here is a selection of fun words Ash learned:</p>

<ul>
  <li>
    <p><strong>Google</strong> (English): Like all my sons, Ash was fascinated by the <a href="https://en.wikipedia.org/wiki/Google_Home">Google
Homes</a> we have everywhere. It will be interesting to see how
their concept of Google evolves over time. To me it is the search engine
company, to them it is the little black speaker that sits on the shelf and
has a personality.</p>
  </li>
  <li>
    <p><strong>Gondola</strong> (English): All three boys love the gondolas at the Oakland Zoo.
Ash learned to say “ganda” very quickly to indicate that he wanted to ride.</p>
  </li>
  <li>
    <p><strong>Cookie</strong> (Spanish): All three boys learned to say cookie in Spanish very
quickly because my parents give them cookies when they visit. Ash actually
tells us that “cookie” in English is wrong and still only calls them
“gagas”.</p>
  </li>
  <li>
    <p><strong>Hulk</strong> (Animal Sounds): Cory loves Hulk and has a large action figure he
plays with. Ash learned from Cory that Hulk makes a roaring sound while
smashing things and would mimic it while playing with the toy.</p>
  </li>
  <li>
    <p><strong>Little Brother</strong> and <strong>Big Brother</strong> (Chinese): Ash learned how to
identify his brothers very quickly, mainly to complain to us when they took
his toys!</p>
  </li>
</ul>

<h2 id="other-writings-on-language-development">Other Writings on Language Development</h2>

<p>If you enjoyed this article, here are all the other articles I wrote about
<a href="/topics/childhood-language/">language development</a>!</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" alt="Black and white photo of two young boys hanging out a window, their faces smudged with soot.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <strong>Comparison of My Three Sons’ Language Development</strong>
    </a>
<br />
I recorded the words my sons spoke as they learned our various languages and now I compare how each developed! Read on to find out how each son learned.  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" alt="Black and white photo of a young boy at a school desk.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <strong>Comparison of My Two Sons’ Language Development</strong>
    </a>
<br />
Being a nerd dad, I recorded all the words my first two sons spoke as they learned them. Now, I compare their language development rate!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <img src="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a man using a blackboard to teach young children punctuation.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <strong>My Second Son’s Language Development</strong>
    </a>
<br />
My second son is a little over two years old. We tracked every word he's spoken to watch his language development, and now you can observe it too!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <img src="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a woman using a blackboard to teach young children how to pronounce words.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <strong>My Son’s Language Development</strong>
    </a>
<br />
My son is a little over two and unfortunately he has two huge nerds for parents. We tracked every word he's spoken to watch his language development, and now you can join us!  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="childhood-language" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[We tracked my third son's language development word by word. Here, in plots, is how he learned to speak. Take a look!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">The Pyramid and the Spire: Management and Individual Contributor Tracks</title>
      <link href="https://alexgude.com/blog/management-vs-ic-the-pyramid-and-the-spire/" rel="alternate" type="text/html" title="The Pyramid and the Spire: Management and Individual Contributor Tracks" />
      <published>2021-11-22T00:00:00-08:00</published>
      <updated>2021-11-22T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/management_vs_ic_the_pyramid_and_the_spire</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/management-vs-ic-the-pyramid-and-the-spire/"><![CDATA[<p>Most tech companies have two tracks for engineers, data scientists, and other
technical people: an individual contributor (IC) track and a manager track.
The intent is that a technical person should be able to advance their career
along the IC track without switching to managing people—which is an entirely
different skill set and job.</p>

<p>I have been on both tracks, and while it is true that you can continue to be
promoted without going into management, there are some trade offs. Many of
these trade offs are discussed in depth elsewhere, but there is one I want to
highlight that is less obvious: <em>There is more room at the top of the
manager track than the IC track.</em></p>

<h2 id="the-structure-of-the-tracks">The Structure of the tracks</h2>

<p>Every company is slightly different, but most of them have their tracks
set up something like this:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>    L1 → L2 → L3 → L4 → L5 → L6 ...
                  ↳ M4 → M5 → M6 ...
</code></pre></div></div>

<p>Where an <strong>L</strong> (short for “Level”) indicates an IC role and an <strong>M</strong> indicates
a manager role. The higher the number, the more senior the role.<sup style="anchor-name:--fnref-numbers" id="fnref:numbers"><a href="#fn:numbers" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>The only role open to the most junior engineers and scientists are IC roles,
which makes sense: junior people lack the hard-earned experience companies
want a manager to have. As they get promoted they move from L1 to L2 and so on
until at some point they reach a level (L4 in my cartoon example above) where
they can move onto the management track (prefixed with M). Of course not every
IC gets this opportunity.</p>

<p>Both tracks narrow as they go up. There are generally a lot (perhaps tens of
thousands for the <a href="https://en.wikipedia.org/wiki/Big_Tech">tech giants</a>) of L1, L2, and L3s in a company, but
L4s and above are rarer. Likewise they have an army of front-line managers
(M4) but only one CEO (or two in rare instances) at the top of the management
track.</p>

<h2 id="the-pyramid-and-the-spire">The Pyramid and the Spire</h2>

<p>The IC track is very wide at the bottom. I was not exaggerating when I said
that there are tens-of-thousands of lower level ICs at the tech giants. But
what is different is how fast they narrow. The IC track gets narrow very
quickly while the manager track stays broad longer. Let me share some examples
from my previous company:</p>

<h3 id="the-highest-levels">The Highest Levels</h3>

<p>At the very highest level, there were a few (probably about five) executive
vice presidents—M9 on my example tracks. There were no L9s at the company.
There was a single L8, someone with more than 30 years of IC experience in
Silicon Valley. The company had more than a dozen M8s.</p>

<h3 id="in-data">In Data</h3>

<p>In the data organization, there was one senior vice presidents (M8), two vice
presidents (M7), and roughly five directors (M6). There were only two L6s, and
one of them had almost 20 years of experience at the company. All of the
managers had less tenure than that.</p>

<h2 id="the-track-forward">The Track Forward</h2>

<p>I am very happy as an IC, I love solving problems and writing code. But I love
helping people, setting strategy, and communicating with partners. I still
haven’t decided which track I will stay on in the long term.</p>

<p>From a pure numbers game, if I want to climb to the highest levels, the
manager track looks better. But it is easier to jump between companies on the
IC track—I could leave my job today and have a stack of comparable IC offers
within a month. Also the manager job is <strong>just different</strong>, I would no longer
be an engineer.</p>

<p>I find <a href="https://twitter.com/mipsytipsy">Charity Majors’s</a> post <a href="https://charity.wtf/2019/01/04/engineering-management-the-pendulum-or-the-ladder/"><em>Engineering Management: The Pendulum Or
The Ladder</em></a> an excellent discussion of the drawbacks of being a manager
and one strategy for jumping back and forth, one that has guided my career
thinking. Give it a look!</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:numbers">

      <p>In this example I started the track at L1, but many companies
start it at a different number. Amazon seems to start at 4, Apple at 2,
Facebook at 3, Google at 3, and Microsoft at 59. For a lot more
information on a specific company’s levels and total compensation, see
<a href="https://www.levels.fyi">levels.fyi</a>. <a href="#fnref:numbers" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
      

      

      
      
        <summary type="html"><![CDATA[Data science has left the era of the Unicorn and entered the era of the team, but that means there is now a whole spectrum of data science jobs. Here is what they do.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/pyramid/the_great_sphinx_david_roberts_1839.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/pyramid/the_great_sphinx_david_roberts_1839.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Data Science, Compensation, and Asking for Money</title>
      <link href="https://alexgude.com/blog/data-science-asking-for-more-money/" rel="alternate" type="text/html" title="Data Science, Compensation, and Asking for Money" />
      <published>2021-10-25T00:00:00-07:00</published>
      <updated>2021-10-25T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_asking_for_more_money</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-asking-for-more-money/"><![CDATA[<hr />

<p><em>This post was meant to be a chapter in a book where data scientists who had
transitioned from academia to industry shared their experience and advice.
Unfortunately, the 2020 pandemic seems to have killed the project, so I am
sharing my chapter here. Enjoy!</em></p>

<hr />

<p>I didn’t become a physicist for money—quite the opposite, in fact. I remember
walking across UC Berkeley’s Sproul Plaza with my father on our way to our
favorite lunch spot when I told him, “I love that <a href="https://en.wikipedia.org/wiki/Cosmology">Cosmology</a> has
so few practical applications! It’s amazing that someone will pay me just to
expand human knowledge!” Perhaps that should have been my first hint that
Physics did not have great long-term prospects, but I don’t think that would
have changed my decision to pursue it.</p>

<p>The first time I visited <a href="https://en.wikipedia.org/wiki/CERN">CERN</a> during <a href="/blog/my-phd-thesis/">my PhD</a>, I got an ascetic
feeling from looking at the dilapidated buildings and missing storm shutters.
Its exterior reflected the same feeling I had expressed to my father a few
years earlier—that science was worth more than dollars. I felt instantly at
home.</p>

<p>When I ended up <a href="/blog/should-i-get-a-phd/">leaving academia</a> several years later, it was for
stability, not money. I had no intention of dragging my family around the
globe just to end up six years older and once again hunting for a job. I’d
seen my CERN colleagues try—and universally fail—to land tenure-track
positions. So I decided to head back to where I grew up—the Bay Area—and see
if I could make it as a data scientist with the help of the <a href="/blog/should-i-go-to-insight/">Insight Data
Science program</a>.</p>

<h2 id="compensation">Compensation</h2>

<p>As someone who’d never been motivated by money, I discovered during my job
search that I was a little naive as to what all the numbers meant.
Fortunately, Insight anticipated this and brought experienced data scientists,
startup founders, and tech company executives to give us a crash course on
life outside of academia. I learned everything there was to know about salary,
bonuses, and equity.</p>

<p>I was already familiar with <strong>base salary</strong>: what you’re paid in semi-weekly
paychecks in exchange for doing your job. It’s what I thought of when I heard
“This job pays $120,000 a year.” But salary isn’t everything—data scientists
often get paid in several different ways.</p>

<p>Every offer I received paid a <strong>yearly bonus</strong> of 15–20% based on performance.
These were from standard tech companies, but I also heard from data scientists
in finance and investment who’d be disappointed if their bonus wasn’t an
“integer multiple of their salary.”<sup style="anchor-name:--fnref-quant_note" id="fnref:quant_note"><a href="#fn:quant_note" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>Insight also spent a lot of time teaching us about <strong>equity</strong>: owning a piece
of the company you’re working in. For <a href="https://en.wikipedia.org/wiki/Public_company">publicly traded companies</a>,
things were simple: I would get a certain number of <a href="https://en.wikipedia.org/wiki/Restricted_stock">RSUs</a> that would
vest over the next three or four years, at which point I could sell them (or
hold them).<sup style="anchor-name:--fnref-rsu_sale" id="fnref:rsu_sale"><a href="#fn:rsu_sale" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>

<p>Startups were more complicated. I didn’t get any offers from startups, but
this is what I learned from Insight and from working with VCs for two years:
particularly valuable early hires can receive up to a few percent ownership of
the startup, but over time that share would <a href="https://en.wikipedia.org/wiki/Stock_dilution">“dilute”</a> as new stock
was issued. Complicated vesting structures could mean that even with a good
exit, your shares might be worth less than you thought if you’re far down the
<a href="https://en.wikipedia.org/wiki/Liquidation_preference">liquidation preference</a> stack.</p>

<p>After all the preparation, I had two offers to consider. But how to choose?
Just pick the biggest number, right? Turns out it isn’t so simple. There are a
lot of <em>intangibles</em>—like “what work would I be doing?” and “who would I
work with?”—that matter a lot, possibly more than money. I decided to go
with the offer that had a bunch of great people on the team and a very
interesting problem, even though the total compensation was lower.</p>

<h2 id="negotiating">Negotiating</h2>

<p>Once I had decided on an offer, I entered the most stressful part:
negotiating. Insight had drilled into me that regardless of the offer I
picked, I had to negotiate it. I spent a few days talking to my friends and
family about the offer, trying to work up the courage to call the hiring
manager. In the end, I took a deep breath and wrote a quick email:</p>

<blockquote>
  <p>Hi!</p>

  <p>I’m very excited by your offer and the thought of working with the team, but
I have another offer for $15,000 more. Can you increase yours? If so, I’ll
sign today.</p>

  <p>Alex</p>
</blockquote>

<p>They wrote back a few hours later, having increased the offer by $7,500. I
signed immediately. That five-minute email earned me $18,000 more over two
years in increased salary and bonus. A pretty good return on investment for a
few minutes of stress.</p>

<p>That time, I had leverage: a higher competing offer. A second offer shows your
true market value—someone else is literally willing to pay you that much.</p>

<p>Two years later, I got an amazing offer from my current company, far higher
than what I’d said I was looking for during the interview. I knew I had made a
slight error giving the first number, and in the back of my mind I could hear
my Insight advisor telling me to negotiate. But I was nervous. I did not have
a counteroffer for leverage or to fall back on.</p>

<p>I talked to the VCs I worked with for a few days to work up the courage. They
all gave the same advice: negotiate. One even said that when he was hired, he
was sure his new boss would have seen it as a red flag if he <em>hadn’t</em> tried to
get a better deal—after all, making deals was his job. Data science is a
little different from investing, but the advice holds: as a hiring manager,
I’ve never once minded when a candidate negotiated. I’m just excited to close
on someone who’s going to help my team reach our goals.</p>

<p>I wrote an email asking for a lot and got a reply back asking for a phone
call. I laid out my rationale—that my friends with similar experience were
getting those kinds of offers. They kept me sweating for a few days, but
finally returned an offer with a $5,000 higher base salary and $50,000 more in
RSUs. It was a lot more stressful than my first negotiating experience because
I did not have a fallback option, but again the return on investment was huge!</p>

<h2 id="helping-others-negotiate">Helping Others Negotiate</h2>

<p>Even if you are not motivated by money, as I was not, if there is some cause
you care about—your family, the environment, helping others—you should
negotiate. Why? Because money will give you the means to advance your aims.
You can donate it, hire someone with it, or put it away for later. In the end,
either you or your employer will have the extra dollars; who do you think will
put it to better use?</p>

<p>One of my friends got a great offer for a software developer position in the
Midwest and was reluctant to negotiate. He asked me, “I have no kids, we own
our house. My wife and I don’t need anything. Should I still negotiate?” I
encouraged him to do so anyway. With just a few emails, he got a larger
signing bonus,<sup style="anchor-name:--fnref-bonus" id="fnref:bonus"><a href="#fn:bonus" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> which, true to his ideals, he donated to charity.</p>

<p>The only leverage he had was that they wanted to hire him and he had not
signed their offer letter yet, but that is often enough. Hiring managers spend
a lot of time screening candidates; when they find one they think is a good
fit, they are anxious to close the deal before someone else does! Offering to
sign immediately if your conditions are met is a great way to open a
negotiation. You signal to the other party that if they do this one thing for
you, they will get what they want: your signature. It’s a win-win.</p>

<p>Another one of my friends got an offer for her first data science job at an
autonomous driving company. She had a few other offers pending, but this was
her dream job. I had helped her through each part of her job search, and so
she called me to talk over the offer. I told her it was a great offer, but
since she still had offers pending, she should ask them for $10,000 more base
salary to sign immediately.</p>

<p>She initially balked; it was her preferred company, after all, and she didn’t
want to lose the offer. I told her they had already spent a lot of money
interviewing her and determined that she was the right candidate. Asking for
more was not going to jeopardize the offer; they would actually appreciate the
opportunity to close on a candidate without having to compete. After a few
minutes, she promised to send them an email. A few days later, she called:
they’d said yes immediately.</p>

<p>Agreeing to sign immediately and having other offers are two good ways to
start your negotiating that I have used and coached my friends in using, but
there are many more. I highly recommend reading <a href="https://www.kalzumeus.com/2012/01/23/salary-negotiation/">Patrick McKenzie’s guide to
negotiating</a>, where he covers everything from what the employer’s
bargaining position is to a step-by-step guide through the whole negotiating
process, and even what to say to avoid the dreaded “So what’s your current
salary?” question. Patrick’s guide has helped engineers and data scientists
earn combined millions of dollars in additional compensation; you should join
their ranks!</p>

<p>Finally, remember that money isn’t everything. You can negotiate all the
various parts of your compensation, including things you might not consider
part of your compensation: vacation days, work location, and the ability to
work from home. These can be easier to negotiate because, as an employee, we
often value them much higher than the exact monetary value assigned by
employers.</p>

<p>Negotiating an offer is stressful but lucrative. Data scientists get many
forms of compensation and are in high demand. That gives us leverage. Use it.
It’s easier than you think, and you truly have nothing to lose.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:quant_note">

      <p>Although as a <a href="https://en.wikipedia.org/wiki/Quantitative_analysis_(finance)">quant</a> friend of mine pointed out: “Zero
is also an integer multiple.” <a href="#fnref:quant_note" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:rsu_sale">

      <p>I generally sell my RSUs immediately; I already have my salary and health
insurance tied up in my employment, so selling the RSUs lets me diversify. <a href="#fnref:rsu_sale" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:bonus">
      <p>A one-time bonus paid when you start a new job. <a href="#fnref:bonus" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Advice about data science salaries and examples from my career of negotiating your offer.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-salaries/Wagner-h%C3%B6henberg_paying_his_dues.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-salaries/Wagner-h%C3%B6henberg_paying_his_dues.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: Differences in Vehicle Collision Rates by Manufacturer During COVID-19</title>
      <link href="https://alexgude.com/blog/switrs-ford-vs-toyota-during-covid-19/" rel="alternate" type="text/html" title="SWITRS: Differences in Vehicle Collision Rates by Manufacturer During COVID-19" />
      <published>2021-09-27T00:00:00-07:00</published>
      <updated>2021-09-27T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/switrs_ford_vs_toyota_during_covid_19</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-ford-vs-toyota-during-covid-19/"><![CDATA[<p>As I prepared to write my post on the <a href="/blog/switrs-covid-19-lockdown-fatal-traffic-collisions/">increase in traffic fatalities during
COVID-19</a>, I made some exploratory plots. One plot made me stop and
stare. Here it is:</p>

<p><a href="/files/switrs-covid/covid_pandemic_ford_vs_toyota_collisions.svg"><img src="/files/switrs-covid/covid_pandemic_ford_vs_toyota_collisions.svg" alt="The number of traffic collisions involving Fords compared to those
involving Toyotas before and after the COVID-19 stay at home order in
California" /></a></p>

<p>This plot isn’t perfect—I will fix it below—but even so it is striking.
Before the <a href="https://en.wikipedia.org/wiki/California_government_response_to_the_COVID-19_pandemic">stay-at-home order</a> the number of collisions involving
<a href="https://en.wikipedia.org/wiki/Toyota">Toyotas</a> was much higher than those involving <a href="https://en.wikipedia.org/wiki/Ford_Motor_Company">Fords</a>. After
the order, the trend flips. Fords have more collisions. I had to figure out
why.</p>

<p>The code for this analysis can be found <a href="/files/switrs-covid/SWITRS%20Ford%20vs%20Toyota%20During%20Lockdown.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-covid/SWITRS%20Ford%20vs%20Toyota%20During%20Lockdown.ipynb">rendered on
Github</a>). The data is available on <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs">Kaggle</a> or
<a href="https://zenodo.org/record/4284843">Zenodo</a>. There is a <a href="https://www.kaggle.com/alexgude/switrs-vehicle-collision-rates-difference-by-make/">hosted Kaggle Notebook</a> version of this
post as well to help you dive right in.</p>

<h2 id="data">Data</h2>

<p>I select all collisions between 2019 and November 30, 2020 that involve a
Toyota or a Ford, with this query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span>
    <span class="p">,</span> <span class="n">p</span><span class="p">.</span><span class="n">vehicle_make</span>
    <span class="p">,</span> <span class="k">count</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">as</span> <span class="n">total</span>
<span class="k">FROM</span> <span class="n">collisions</span> <span class="k">AS</span> <span class="k">c</span>
<span class="k">LEFT</span> <span class="k">JOIN</span> <span class="n">parties</span> <span class="k">as</span> <span class="n">p</span>
<span class="k">ON</span> <span class="n">p</span><span class="p">.</span><span class="n">case_id</span> <span class="o">=</span> <span class="k">c</span><span class="p">.</span><span class="n">case_id</span>
<span class="k">WHERE</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span> <span class="k">BETWEEN</span> <span class="s1">'2019-01-01'</span> <span class="k">AND</span> <span class="s1">'2020-11-30'</span>
<span class="k">AND</span> <span class="n">p</span><span class="p">.</span><span class="n">vehicle_make</span> <span class="k">IN</span> <span class="p">(</span><span class="s1">'ford'</span><span class="p">,</span> <span class="s1">'toyota'</span><span class="p">)</span>
<span class="k">GROUP</span> <span class="k">BY</span> <span class="mi">1</span><span class="p">,</span> <span class="mi">2</span><span class="p">;</span>
</code></pre></div></div>

<p>I start the data in 2019 because I need a sample from <em>before</em> the pandemic
changed behavior, but I didn’t want to go too far back because <a href="/blog/switrs-crashes-by-date/#crashes-per-week">collision
rates vary drastically year-to-year</a>. I cut off the data in
November because the reporting is not yet complete for December.</p>

<h2 id="normalized-collision-rate">Normalized Collision Rate</h2>

<p>The number of collisions depends on many factors, primary among them is
<a href="https://en.wikipedia.org/wiki/Units_of_transportation_measurement#Fatalities_by_VMT">vehicle miles traveled</a>.<sup style="anchor-name:--fnref-mn_pub_safety" id="fnref:mn_pub_safety"><a href="#fn:mn_pub_safety" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> To help control for VMT, I
normalize the mean number of collisions for each make of vehicle from January
through June of 2019. This gives me a baseline to compare against. Here is the
normalized plot:</p>

<p><a href="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions.svg"><img src="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions.svg" alt="The collision rate for Fords compared to Toyotas before and after the COVID-19 stay at home order in
California, with mean normalized from January 2019 through June 2019." /></a></p>

<h2 id="interpretation">Interpretation</h2>

<p>The normalized rates match up well through the <a href="/blog/switrs-crashes-by-date//#day-by-day">Christmas and New Year
holidays</a>, which is the two-week dip caused by people taking time off
work and hence not commuting. But right after, the series diverge:</p>

<ul>
  <li>
    <p>Toyota collisions trend down a few weeks before the stay-at-home order and
drop off significantly the week before. Ford collisions stay constant until
the order. This suggests that Toyota drivers made a decision to stay home by
themselves while Ford drivers waited until the state mandated it.</p>
  </li>
  <li>
    <p>Toyota collisions drop much more the week of the order, likely indicating
that more Toyota drivers stayed at home when told to do so.</p>
  </li>
  <li>
    <p>Ford collisions recovered towards their pre-pandemic level faster than
Toyota collisions, indicating that Ford drivers were quicker to get back on
the road when allowed to.</p>
  </li>
</ul>

<p>Taken together, I think these observations suggest the difference is due to a
<a href="https://en.wikipedia.org/wiki/White-collar_worker">white-collar</a>–<a href="https://en.wikipedia.org/wiki/Blue-collar_worker">blue-collar</a> divide. White-collar
workers generally have more flexible work arrangements and their jobs are
easier to do from home, whereas blue-collar workers have to travel to a job
site to perform their work. Blue-collar workers are more conservative than
white-collar workers and more likely to buy American branded cars like
Fords.<sup style="anchor-name:--fnref-political_cars" id="fnref:political_cars"><a href="#fn:political_cars" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>

<p>Initially I thought this difference would be driven purely by the prevalence
of Ford trucks, but as we shall see it is not just trucks versus cars.</p>

<h3 id="trucks">Trucks</h3>

<p>Is it just that there are more Ford trucks? No.</p>

<p><a href="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_trucks.svg"><img src="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_trucks.svg" alt="The collision rate for Ford trucks compared to Toyota trucks before and
after the COVID-19 stay at home order in California, with mean normalized from
January 2019 through June 2019." /></a></p>

<p>The same pattern holds, although both makes recover faster, with Fords
returning to pre-pandemic levels and Toyota getting to 80%, which is much
higher than the 50% Toyota reached when including non-trucks.</p>

<h3 id="location">Location</h3>

<p>Perhaps Ford owners just live in areas with looser restrictions, like the
Central Valley? No. Here is data from Contra Costa County, part of the Bay
Area:</p>

<p><a href="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_contra_costa.svg"><img src="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_contra_costa.svg" alt="The collision rate for Fords compared to Toyotas in Contra Costa County
before and after the COVID-19 stay at home order in California, with mean
normalized from January 2019 through June
2019." /></a></p>

<p>It is the same pattern, but with a lot more noise due to the smaller
population.</p>

<h3 id="age">Age</h3>

<p>Young drivers get in more accidents. Perhaps there is a strong age difference
driving the trend? There is an age difference, see:</p>

<p><a href="/files/switrs-covid/covid_pandemic_ford_vs_toyota_collisions_age_distribution.svg"><img src="/files/switrs-covid/covid_pandemic_ford_vs_toyota_collisions_age_distribution.svg" alt="Area normalized distribution of Toyota and Ford driver ages during the
COVID-19 stay at home order in California." /></a></p>

<p>But that alone doesn’t account for the pattern:</p>

<p><a href="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_age_30_to_50.svg"><img src="/files/switrs-covid/covid_pandemic_normalized_ford_vs_toyota_collisions_age_30_to_50.svg" alt="The collision rate for Fords compared to Toyotas for drivers aged 30 to 50
before and after the COVID-19 stay at home order in California, with mean
normalized from January 2019 through June
2019." /></a></p>

<h3 id="putting-it-all-together">Putting It All Together</h3>

<p>A person’s identity is made up of many traits: their age, their politics,
where they live, what job they do, and yes, what car they drive. I looked at
three different traits—vehicle type, location, and age—and none of them
explain the entirety of the collision rate difference between Toyotas and
Fords after the COVID-19 stay-at-home order. My conclusion is that Ford
drivers are just different from Toyota drivers, in multiple ways, each of
which contributes to the trend.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:mn_pub_safety">
      <p>From the Minnesota Department of Public Safety:</p>

      <p><figure class="cited-quote"><blockquote cite="https://web.archive.org/web/20151004145329/https://dps.mn.gov/divisions/ots/reports-statistics/Documents/2014-crash-facts.pdf"> <p>Volume of traffic, or vehicle miles traveled (VMT), is a predictor of crash incidence. All other things being equal, as VMT increases, so will traffic crashes. The relationship may not be simple, however; after a point, increasing congestion leads to reduced speeds, changing the proportion of crashes that occur at different severity levels.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Minnesota Department of Public Safety, Office of Traffic Safety. <a href="https://web.archive.org/web/20151004145329/https://dps.mn.gov/divisions/ots/reports-statistics/Documents/2014-crash-facts.pdf"><cite>Minnesota Traffic Crashes in 2014</cite></a>. 2014. pp. 2.</span></figcaption></figure> <a href="#fnref:mn_pub_safety" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:political_cars">

      <p>The type of car and brand both are driven by political leaning:</p>

      <p><figure class="cited-quote"><blockquote cite="https://www.nytimes.com/2005/04/01/automobiles/your-car-politics-on-wheels.html"> <p>The most left-leaning models with at least a dozen sightings in Mr. MacMichael’s project were the Honda Civic (80-20 left-leaning), Toyota Corolla (78-19) and Toyota Camry (74-26). The list of most right-leaning was led by another Toyota, but a midsize SUV, the Toyota 4Runner (86-14), followed by the Ford Expedition (76-24) and Ford F-150 (75-25).</p>  </blockquote><figcaption>—<span markdown="0" class="citation">Tierney, John. <a href="https://www.nytimes.com/2005/04/01/automobiles/your-car-politics-on-wheels.html">“Your Car: Politics on Wheels”</a> <cite>The New York Times</cite>. April 1, 2005.</span></figcaption></figure> <a href="#fnref:political_cars" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[California was put under a stay-at-home order in March, 2020. Toyota drivers stayed home, Ford drivers did not; why?!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-covid/mail_truck_tries_to_climb_tree_in_boston_1927.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-covid/mail_truck_tries_to_climb_tree_in_boston_1927.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Plotting the 2021 Tour de France</title>
      <link href="https://alexgude.com/blog/2021-tour-de-france-plot/" rel="alternate" type="text/html" title="Plotting the 2021 Tour de France" />
      <published>2021-08-18T00:00:00-07:00</published>
      <updated>2021-08-18T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/2021_tour_de_france_plot</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/2021-tour-de-france-plot/"><![CDATA[<p>The 108th edition of the <a href="https://en.wikipedia.org/wiki/2021_Tour_de_France">Tour de France</a> started in late June this
year. The start date was shifted back slightly to avoid overlapping with the
<a href="https://en.wikipedia.org/wiki/2020_Summer_Olympics">rescheduled 2020 Summer Olympics</a>, but the race was otherwise
unaffected by the <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">ongoing COVID pandemic</a> which had forced the
postponement of last year’s race. In this post, just like <a href="/blog/2020-tour-de-france-plot/">last
year’s</a>, I plot how the race unfolded.</p>

<p>The code that generated the plots can be found <a href="/files/tour-de-france//Tour%20de%20France%202021%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/tour-de-france//Tour%20de%20France%202021%20Plot.ipynb">rendered on Github</a>). The data <a href="/files/tour-de-france//2021-tdf-dataframe.json">is here</a>.</p>

<h2 id="the-race-for-yellow">The Race for Yellow</h2>

<p>The top award in the Tour is the <a href="https://en.wikipedia.org/wiki/General_classification_in_the_Tour_de_France">yellow jersey</a>, which is awarded to
the rider with the lowest combined time across the 21 stages of the race.
<a href="https://en.wikipedia.org/wiki/Tadej_Poga%C4%8Dar">Tadej Pogačar</a>, the incredibly young<sup style="anchor-name:--fnref-young" id="fnref:young"><a href="#fn:young" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and surprisingly dominant
winner of last year’s race, was the clear favorite this year.</p>

<p><a href="https://en.wikipedia.org/wiki/Primo%C5%BE_Rogli%C4%8D">Primož Roglič</a> was a favorite again after <a href="/blog/2020-tour-de-france-plot/">taking second place in
last year’s Tour</a>. Since then he had defended his title in one of
the two other grand tours, the <a href="https://en.wikipedia.org/wiki/2020_Vuelta_a_Espa%C3%B1a">Vuelta a España</a>, winning for a second
year in a row.</p>

<p><a href="https://en.wikipedia.org/wiki/Ineos_Grenadiers">Ineos Grenadiers</a> teammates <a href="https://en.wikipedia.org/wiki/Richie_Porte">Richie Porte</a>, <a href="https://en.wikipedia.org/wiki/Geraint_Thomas">Geraint
Thomas</a>, and <a href="https://en.wikipedia.org/wiki/Richard_Carapaz">Richard Carapaz</a> were also in the running.
Porte had won the <a href="https://en.wikipedia.org/wiki/2021_Crit%C3%A9rium_du_Dauphin%C3%A9">Critérium du Dauphiné</a>—considered a warm-up race for
the Tour used to test a rider’s form—and Thomas had placed third in the
Critérium and was the 2018 Tour winner. Carapaz had just won the <a href="https://en.wikipedia.org/wiki/2021_Tour_de_Suisse">Tour de
Suisse</a>, the other Tour warm-up race, and had won the <a href="https://en.wikipedia.org/wiki/Giro_d%27Italia">Giro
d’Italia</a> in 2019.</p>

<p>Unfortunately, the race for yellow turned out to be far less exciting than
last year, indicated by the large time gap between Pogačar and the rest of the
field that formed early in the race:</p>

<p><a href="/files/tour-de-france//2021_tour_de_france_top_5.svg"><img src="/files/tour-de-france//2021_tour_de_france_top_5.svg" alt="A line plot showing how far behind the leader each top-finishing rider was
after each stage of the 2021 Tour de France." /></a></p>

<p>Pogačar only faltered on stage 7 as the race entered the Alps. He dropped to
four minutes behind current race leader <a href="https://en.wikipedia.org/wiki/Mathieu_van_der_Poel">Mathieu van der Poel</a> and three
and half minutes behind second place <a href="https://en.wikipedia.org/wiki/Wout_van_Aert">Wout van Aert</a>.<sup style="anchor-name:--fnref-cyclocross" id="fnref:cyclocross"><a href="#fn:cyclocross" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> But
Pogačar stormed back on stage 8 to take the lead, which he maintained for the
rest of the race.</p>

<p><a href="https://en.wikipedia.org/wiki/Ben_O%27Connor_(cyclist)">Ben O’Connor</a> attempted to contest Pogačar’s lead on stage 9 with a
solo break away win but fell two minutes short. O’Connor’s effort put him in
second place, but he wasn’t able to hold his form and eventually placed fourth
after losing time in the Pyrenees. By the end of stage 10, Pogačar had an
unassailable lead of almost 6 minutes.</p>

<h3 id="the-green-jersey">The Green Jersey</h3>

<p>Although the race for yellow was uneventful, the race for the <a href="https://en.wikipedia.org/wiki/Points_classification_in_the_Tour_de_France">green
jersey</a> was incredibly exciting. The green jersey is awarded to the
rider with the most points, which are earned by winning intermediate sprints
and stages.</p>

<p><a href="https://en.wikipedia.org/wiki/Mark_Cavendish">Mark Cavendish</a>—considered by some to be the greatest sprinter
ever<sup style="anchor-name:--fnref-nostalgia" id="fnref:nostalgia"><a href="#fn:nostalgia" class="footnote" rel="footnote" role="doc-noteref">3</a></sup>—entered the race having won 30 stages of the Tour, just
<a href="https://en.wikipedia.org/wiki/Tour_de_France_records_and_statistics#Stage_wins_per_rider">four behind</a> all-time great <a href="https://en.wikipedia.org/wiki/Eddy_Merckx">Eddy Merckx</a>. But
Cavendish had not won a Tour sprint since 2016 or any stage of any race since
2018. His performance had fallen so far that he had considered retiring before
the 2021 season. But in early 2021, he showed a return to his winning
form with four dominate sprint wins in the <a href="https://en.wikipedia.org/wiki/2021_Presidential_Tour_of_Turkey">Tour of Turkey</a>, which
raised the possibility of him beating Merckx’s record.</p>

<p>Here is how the sprint race turned out, with sprint stages shaded in grey:</p>

<p><a href="/files/tour-de-france//2021_tour_de_france_top_5_sprint.svg"><img src="/files/tour-de-france//2021_tour_de_france_top_5_sprint.svg" alt="A line plot showing how far behind the points leader the top five sprint
sprinters were." /></a></p>

<p>Cavendish took the lead in the points competition with a win on stage 4. He
extended his lead on stage 6 with another win which brought him up to 32
all-time, just 2 behind Merckx. Cavendish had to survive the Alps in stages 7
and 8 if he wanted another shot at sprint wins. He managed to avoid the time
cuts with the help of his team and went on to win twice more on stages 10 and
13, where second place <a href="https://en.wikipedia.org/wiki/Michael_Matthews_(cyclist)">Michael Matthews</a> falls to his lowest point
before clawing his way back over the next few stages.</p>

<p>Cavendish’s win on stage 13 tied Merckx’s record of 34 tour wins and set him
up to beat the record during the <a href="https://en.wikipedia.org/wiki/Champs-%C3%89lys%C3%A9es_stage_in_the_Tour_de_France">final sprint of the tour on the
Champs-Élysées</a>. Unfortunately it was not to be, Cavendish came in
third on the final sprint behind Wout van Aert and <a href="https://en.wikipedia.org/wiki/Jasper_Philipsen">Jasper
Philipsen</a>. Cavendish may have another chance to beat the record in
2022, but as he is at the tail-end of his career it is not certain he will
make the race.</p>

<h2 id="the-rest-of-the-race">The Rest of the Race</h2>

<p>184 riders started the race and 141 finished. Here is how each rider fared:</p>

<p><a href="/files/tour-de-france//2021_tour_de_france.svg"><img src="/files/tour-de-france//2021_tour_de_france.svg" alt="A line plot showing how far behind the leader every rider was for each
stage." /></a></p>

<p>Cavendish paced himself in the mountains, finishing with the slowest rider to
save his energy. Cavendish’s teammate, <a href="https://en.wikipedia.org/wiki/Tim_Declercq">Tim Declercq</a>’s time is
almost identical to Cavendish’s for the first 12 stages, as he stayed with
the sprinter to ensure that Cavendish made it in under the time cut. Declercq
was involved in a major crash in stage 13, where he lost almost 15 minutes.
But he held on as other, slower riders dropped out, allowing him take the
<a href="https://en.wikipedia.org/wiki/Lanterne_rouge">lanterne rougue</a> awarded to the last place rider.</p>

<p>Finally, how did the other riders who held the yellow jersey during the race,
<a href="https://en.wikipedia.org/wiki/Julian_Alaphilippe">Julian Alaphilippe</a> and Mathieu van der Poel, do? Van der Poel
dropped out when they hit the mountains to prepare for the Olympics. Despite
his strength, he is not the type of rider who could have won this tour, being
too heavy to climb quickly. Alaphilippe, although not a climber specialist, can
compete in the mountains as <a href="/blog/2019-tour-de-france-plot/#the-race-for-yellow">we saw in the 2019 Tour</a>, but he too
started to lose time in the Alps. He held on and finished in Paris, but far
down the leaderboard.</p>

<p>Although this year’s competition for the yellow jersey lacked the excitement
of last year’s Tour, Mark Cavendish’s amazing return to form provided some
rare sprinting tension. Hopefully Cavendish will return next year to attempt
to break Merckx’s record!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:young">

      <p>Pogačar is the second youngest winner of the Tour at 21. <a href="https://en.wikipedia.org/wiki/Henri_Cornet">Henri
Cornet</a> is the youngest winner at just 10 days shy of 20 when he
won the 1904 edition. <a href="#fnref:young" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:cyclocross">

      <p>Van der Poel and van Aert are an exciting pair to watch! They got their
start dominating <a href="https://en.wikipedia.org/wiki/Cyclo-cross">cyclo-cross</a>, where their only real competition
was each other. Van der Poel has won four of the last seven Cyclo-cross World
Championships, and van Aert has won the other three. Both have continued their
domination—and rivalry—on the road in the last few years. <a href="#fnref:cyclocross" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:nostalgia">

      <p>Mark Cavendish, or just “Cav” to the fans, is a rider to whom I feel a strong
connection. He was the sprinter to beat when I first started watching
cycling and was one of the few riders I was able to recognize and watch
for in the race. I learned the tactics and tricks of sprinting from
watching Cav and his leadout train.</p>

      <p>The end of his dominance came at roughly the same time that my interest in
cycling started to wane. Most of the riders I had come into the sport with
were leaving the peloton, the teams had all be renamed and many had broken
up, and my life was becoming busier making following the sport hard.</p>

      <p>But with Cav returning to his old dominance of the sprint this year, I
felt like I was back in 2013 watching cycling for the first time. It gave
me a sense of nostalgia and excitement for the sport I hadn’t felt for a
while. I hope the feeling lasts. <a href="#fnref:nostalgia" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[The 2021 Tour de France turned out much differently from last year's edition! See exactly how it unfolded in this post.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_swiss_team.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_swiss_team.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: Increase In Traffic Fatalities After COVID-19 Lock Down</title>
      <link href="https://alexgude.com/blog/switrs-covid-19-lockdown-fatal-traffic-collisions/" rel="alternate" type="text/html" title="SWITRS: Increase In Traffic Fatalities After COVID-19 Lock Down" />
      <published>2021-07-19T00:00:00-07:00</published>
      <updated>2021-07-19T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/switrs_covid_19_lockdown_fatal_traffic_collisions</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-covid-19-lockdown-fatal-traffic-collisions/"><![CDATA[<p>California had its <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic_in_California">first case of COVID-19</a> on January 26, 2020. The
Governor mandated a <a href="https://en.wikipedia.org/wiki/California_government_response_to_the_COVID-19_pandemic">state-wide stay-at-home order</a> on March 19, 2020.
The morning and evening commutes stopped immediately. Traffic volume decreased
by more than 50% and stayed low for weeks. Slowly the restrictions were
relaxed and traffic returned, but has still not reached pre-pandemic levels.</p>

<p>The number of traffic collisions <strong>decreased</strong> as you would expect with the
decreased volume but, surprisingly, the severity of the collisions
<strong>increased</strong>. The <a href="https://www.nhtsa.gov/press-releases/2020-fatality-data-show-increased-traffic-fatalities-during-pandemic">rate of fatal accidents increased across the
country</a>. The National Highway Traffic Safety Administration attributes
the increase to a change in behavior by drivers who stayed on the road: they
drove more recklessly and wore their seatbelt less often.<sup style="anchor-name:--fnref-nhtsa" id="fnref:nhtsa"><a href="#fn:nhtsa" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>I can’t test that hypothesis with my <a href="/blog/switrs-sqlite-hosted-dataset/">SWITRS data</a>—it
does not include much information about driving behavior, only about
collisions—but I can look at the fatality rate on California roads.</p>

<p>The code for this analysis can be found <a href="/files/switrs-covid/SWITRS%20Fatalities%20During%20COVID%20Lockdown.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-covid/SWITRS%20Fatalities%20During%20COVID%20Lockdown.ipynb">rendered on
Github</a>). The data is available on <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs">Kaggle</a> or
<a href="https://zenodo.org/record/4284843">Zenodo</a>. There is a <a href="https://www.kaggle.com/alexgude/switrs-increase-in-traffic-fatalities-after-covid">hosted Kaggle Notebook</a> version of this
post as well to help you dive right in.</p>

<h2 id="data">Data</h2>

<p>I selected all collisions in the dataset between the start of 2019 and
November 30th, 2020, including whether there was a fatality as a result of the
collision, with this query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">collision_date</span>
    <span class="p">,</span> <span class="mi">1</span> <span class="k">AS</span> <span class="n">crashes</span>
    <span class="p">,</span> <span class="n">IIF</span><span class="p">(</span><span class="n">collision_severity</span><span class="o">=</span><span class="s1">'fatal'</span><span class="p">,</span> <span class="mi">1</span><span class="p">,</span> <span class="mi">0</span><span class="p">)</span> <span class="k">AS</span> <span class="n">fatalities</span>
<span class="k">FROM</span> <span class="n">collisions</span>
<span class="k">WHERE</span> <span class="n">collision_date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">collision_date</span> <span class="k">BETWEEN</span> <span class="s1">'2019-01-01'</span> <span class="k">AND</span> <span class="s1">'2020-11-30'</span>
</code></pre></div></div>

<p>I start the data in 2019 because I need a sample from <em>before</em> the pandemic
changed behavior, but I didn’t want to go too far back because <a href="/blog/switrs-crashes-by-date//#crashes-per-week">collision
rates vary drastically year-to-year</a>. I cut off the data in
November because the reporting is not yet complete for December.</p>

<h2 id="fatality-rate">Fatality Rate</h2>

<p>I calculate the weekly fatality rate. It is the number of traffic collisions
that resulted in a fatality divided by the total number of collisions during
the week. Here is what that rate looks like before and after the stay-at-home
order:</p>

<p><a href="/files/switrs-covid/fatality_rate_per_week_in_california_after_covid.svg"><img src="/files/switrs-covid/fatality_rate_per_week_in_california_after_covid.svg" alt="The traffic fatality rate in California before and after the COVID-19
stay-at-home order." /></a></p>

<p>You can see the fatality rate <strong>immediately</strong> jumps up to over 1% for the
first time in our dataset, and then goes even higher in the coming weeks. It
stays elevated for the entirety of our data range.</p>

<p>Another way to look at this data is to plot a histogram of the rate before
and after the stay-at-home order. Here it is:</p>

<p><a href="/files/switrs-covid/fatality_rate_per_week_in_california_after_covid_histograms.svg"><img src="/files/switrs-covid/fatality_rate_per_week_in_california_after_covid_histograms.svg" alt="A histogram showing traffic fatality rate in California before and after
the COVID-19 stay-at-home order." /></a></p>

<p>The weeks with the highest fatality rate before the pandemic are between 0.8%
and 1%. These overlap with the <em>lowest</em> fatality rate weeks after the
stay-at-home order. These are clearly different distributions, but we can
quantify that difference with a <a href="https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney_U_test">Mann–Whitney <em>U</em> test</a>.</p>

<p>The Mann–Whitney test compares the probability that a value randomly drawn
from the first distribution is larger than one randomly drawn from the second
distribution, with a correction for ties. If this probability is not 50% (as
it would be if they were the same) then the distributions must be different.
The test is nonparametric and only assumes that the observations are
independent, that they are orderable, that under the null hypothesis the
distributions are equal, and under the alternative hypothesis the
distributions are different.</p>

<p>The test confirms our eye test with a <em>p</em>-value of 3.6e-14. These
distributions are significantly different, meaning that the California
stay-at-home order increased the traffic fatality rate.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:nhtsa">
      <p>Specifically:</p>

      <p><figure class="cited-quote"><blockquote cite="https://www.nhtsa.gov/press-releases/2020-fatality-data-show-increased-traffic-fatalities-during-pandemic"> <p>NHTSA’s research suggests that throughout the national public health emergency and associated lockdowns, driving patterns and behaviors changed significantly, and that drivers who remained on the roads engaged in more risky behavior, including speeding, failing to wear seat belts, and driving under the influence of drugs or alcohol. Traffic data indicates that average speeds increased throughout the year, and examples of extreme speeds became more common, while the evidence also shows that fewer people involved in crashes used their seat belts.</p>  </blockquote><figcaption>—<span markdown="0" class="citation">National Highway Traffic Safety Administration. <a href="https://www.nhtsa.gov/press-releases/2020-fatality-data-show-increased-traffic-fatalities-during-pandemic"><cite>2020 Fatality Data Show Increased Traffic Fatalities During Pandemic</cite></a>. June 3, 2021.</span></figcaption></figure> <a href="#fnref:nhtsa" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[California was put under a stay-at-home order in March, 2020. As expected, traffic volume decreased, but what happened to rate of fatal accidents? They skyrocketed!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-covid/auto_accident_on_bloor_street_west_in_1918.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-covid/auto_accident_on_bloor_street_west_in_1918.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">The Data Science Spectrum: From Analyst to Machine Learning</title>
      <link href="https://alexgude.com/blog/data-science-job-spectrum/" rel="alternate" type="text/html" title="The Data Science Spectrum: From Analyst to Machine Learning" />
      <published>2021-06-01T00:00:00-07:00</published>
      <updated>2021-06-01T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_job_spectrum</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-job-spectrum/"><![CDATA[<p>The role of a data scientist has become narrower and more specialized as the
demand for them has increased. In my last post, <a href="/blog/the-data-science-split/"><em>The Data Science Split</em></a>,
I talked about why I think this happened. In this post, I will walk through a
few of the most common roles in the data ecosystem and cover what they do and
what their skill sets are.</p>

<p>It is useful to know where you prefer to be on the data science spectrum, as it
will determine what roles you should apply for. “What are the responsibilities
of this position and what the key skills to be successful in it?” is one of
the first questions I ask when applying to a new position. The answer lets me
map that specific role to its place in the ecosystem and helps me determine if
I would be interested in the job.</p>

<h2 id="the-spectrum">The Spectrum</h2>

<p>You could define the spectrum of data science along multiple axes, but I
find using just one works pretty well:<sup style="anchor-name:--fnref-research" id="fnref:research"><a href="#fn:research" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p><strong>Engineeriness</strong>: Roughly, how close the role is to a traditional software
engineering role.</p>

<p>On the “low engineeriness” side of the spectrum you have roles that work
almost entirely with the contents of the data and domain-specific languages
for data access, processing, and plotting. As you move towards the other end
you start working with “lower-level” languages and often less on the
data itself and more on supportive tooling around it.</p>

<p>A particular job at a company could fall anywhere on the spectrum, but the
title gives a good idea of where exactly it fits. Below are five common job
titles in a rough order from least to most “engineery”. Of course, in the real
world, the responsibilities of these jobs overlaps heavily with their
neighbors on the spectrum.</p>

<h3 id="business-analyst">Business Analyst</h3>

<p>A <em>business analyst</em><sup style="anchor-name:--fnref-biz" id="fnref:biz"><a href="#fn:biz" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> uses data to help the company understand what has
happened, what is happening, and what will likely happen so they can make
better decisions. Their primary deliverables are internal-facing reports,
dashboards, and presentations. They are generally really adept at SQL and
making data visualizations, but are less likely to use general-purpose
languages like Python.</p>

<h3 id="data-scientist">Data Scientist</h3>

<p>A <em>data scientist</em><sup style="anchor-name:--fnref-ds" id="fnref:ds"><a href="#fn:ds" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> is an expert at statistics and experimental design.
They don’t just plot trends, they understand what causes them, and how you can
influence them. They can clean a dataset, find biases, and then use it to
power decisions and products. They work with more general programming
languages like R or Python.</p>

<h3 id="modeler">Modeler</h3>

<p><em>Machine learning modeler</em><sup style="anchor-name:--fnref-mlm" id="fnref:mlm"><a href="#fn:mlm" class="footnote" rel="footnote" role="doc-noteref">4</a></sup> is a rarer title, but I include it because I
feel it fills the hole between data scientist and machine learning engineer.
This role focuses on building models that directly impact customers. They find
customer problems, build machine learning models to solve them, and own those
models all the way from first iteration through hosting it in production. They
use Python, machine learning frameworks like TensorFlow, sometimes Scala and
Spark, Docker, and REST APIs.</p>

<p>This is the role I feel most comfortable in, with its mix of <a href="/topics/software-development/">software
development</a>, <a href="/topics/machine-learning/">machine learning</a>, and direct impact on customers.</p>

<h3 id="machine-learning-engineer">Machine Learning Engineer</h3>

<p>A <em>machine learning engineer</em><sup style="anchor-name:--fnref-mle" id="fnref:mle"><a href="#fn:mle" class="footnote" rel="footnote" role="doc-noteref">5</a></sup> focuses on the platforms underlying
machine learning modeling and hosting. They often build ML tooling, hosting,
and pieces of ML specific infrastructure like feature stores. They focus on
making sure the machine learning models can scale to meet the demands of
running in production and return answers fast enough to be used. They
generally work with lower-level languages than the modelers like Scala or
Java. Many MLEs come from a software engineering background.</p>

<h3 id="data-engineer">Data Engineer</h3>

<p><em>Data engineers</em> build the infrastructure the data flows through. All the SQL
databases, NoSQL, queues, streams, etc. that power the business and allow the
other data roles to make use of it. They’re experts in a cloud services (where
these systems are mostly hosted) and scaling systems to meet the demands of
millions or billions of users while collecting and organizing their data.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:research">

      <p>If I were to add a second axis, it would probably be
<strong>Researchiness</strong> to differentiate the product focused data roles covered
in this post from the more academic roles present at some large companies.
The biggest difference is “publishing papers” is a metric more researchy
roles track. <a href="#fnref:research" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:biz">

      <p>Sometimes data analyst, business intelligence analyst, or even data
scientist. <a href="#fnref:biz" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ds">

      <p>Also product data scientist, sometimes decision scientist,
statistician. These align closely with <a href="https://twitter.com/michaelhochster">Michael
Hochster’s</a> <a href="https://www.quora.com/What-is-data-science/answer/Michael-Hochster">Type A Data Scientists</a>. <a href="#fnref:ds" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:mlm">

      <p>This role is sometimes called data scientist, sometimes machine
learning engineer; often those two roles split the responsibility. These
roles are closer to <a href="https://twitter.com/michaelhochster">Michael Hochster’s</a> <a href="https://www.quora.com/What-is-data-science/answer/Michael-Hochster">Type B Data
Scientists</a>. <a href="#fnref:mlm" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:mle">
      <p>Sometime machine learning infrastructure engineer. <a href="#fnref:mle" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="data-science" />
        
          <category term="interviewing" />
        
      

      

      
      
        <summary type="html"><![CDATA[Data science has left the era of the Unicorn and entered the era of the team, but that means there is now a whole spectrum of data science jobs. Here is what they do.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-spectrum/tribeam_prism.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-spectrum/tribeam_prism.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">The Data Science Split: From Unicorns to Teams</title>
      <link href="https://alexgude.com/blog/the-data-science-split/" rel="alternate" type="text/html" title="The Data Science Split: From Unicorns to Teams" />
      <published>2021-05-31T00:00:00-07:00</published>
      <updated>2021-05-31T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/the_data_science_split</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/the-data-science-split/"><![CDATA[<p>I discovered data science as a possible career in 2014 when I was looking for
something to do after <a href="/blog/should-i-get-a-phd/#but-there-are-no-jobs">deciding not to pursue a career in academia</a>. Less
than a year later I was a professional data scientist,<sup style="anchor-name:--fnref-pro" id="fnref:pro"><a href="#fn:pro" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> having moved
across country and gotten a job with the help of <a href="/blog/should-i-go-to-insight/">Insight</a>. It was
there that I realized that the data science role was already splintering into
multiple, specialized roles.</p>

<h2 id="the-split">The Split</h2>

<p>At Insight I ran into lots of different types of data scientists. Some were
really, really good at statistics; others lived for experiment design and
power analysis; some—like me—loved writing clean, performant code and
playing around with deep learning. Yet at the end of the program we mostly
interviewed for jobs that were no more specific than “data scientist”. The
companies were looking for someone to set up their data teams and extract
value from their data.</p>

<p>This was the tail end of the unicorn era. Data engineering had already become
its own specialty, but everything else was still under the umbrella term data
scientist. Even as a neophyte I could see further splits coming.</p>

<h3 id="unicorn-era">Unicorn Era</h3>

<p>The <strong>Unicorn Era</strong> was the time when data science was first being
established. Companies had heard that “data was the new oil” and were
desperate to hire someone to refine it for them. They looked to hire a person
who could span the entire spectrum of data jobs, from setting up the
infrastructure, to building analyses on top of it, to shipping and owning
models in production. But people who could master all of these skills were not
just rare, they were mythical. Hence, unicorns.</p>

<p>It worked OK for awhile. There were few data science jobs so companies needed
to find only a scant handful of unicorns. Still it couldn’t last. Demand for
data scientists exploded and the tasks they were asked to perform became more
and more demanding and more and more specialized. It no longer made sense to
try to find a single person to do it all, for many reasons.</p>

<p>First, finding unicorns had always been nearly impossible. Every data
scientist is an expert in some part of the field, but almost no one is great
at everything <strong>and</strong> interested in all of it. Splitting the role allowed
people to focus on the areas that they were most passionate about.</p>

<p>Second, training a data scientist—waiting for them to finish a PhD in some
<em>other</em> subject first while hoping the required data skills would rub off on
them in the process—took <strong>too long</strong>. An undergraduate degree was the
obvious solution and every college quickly spun up a data science
undergraduate program. But how could you impart the ten years of knowledge
gathered in the lead up to a PhD in just four? You can’t— unless you split
it across three different people.</p>

<h3 id="the-team-era">The Team Era</h3>

<p>The data science role slowly split into multiple positions. Work that had
before been owned by individuals with ill defined titles instead was
distributed across an entire teams. Hence, the <strong>Team Era</strong>.</p>

<p>But it didn’t split into just one or two new roles, it split into a whole
spectrum. The roles cover everything from setting up the underlying
infrastructure, to improving specialized tooling, to building models, to
reporting out the results. I discuss five of these roles in my next post:</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/data-science-job-spectrum/">
      <img src="https://alexgude.com/files/data-science-spectrum/tribeam_prism.jpg" alt="A triangular prism breaking white light into its components.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/data-science-job-spectrum/">
      <strong>The Data Science Spectrum: <br />From Analyst to Machine Learning</strong>
    </a>
<br />
Data science has left the era of the Unicorn and entered the era of the team, but that means there is now a whole spectrum of data science jobs. Here is what they do.  </div>
</li>
</ul>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:pro">

      <p>At least in the sense that I was paid to do it… I don’t claim to
have been any good. <a href="#fnref:pro" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[When data science started the job covered everything from setting up databases to running experiments to making models. But finding Unicorns was impossible; something had to give.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-spectrum/eugene_f_kranz_1965.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-spectrum/eugene_f_kranz_1965.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Where to Host Public Datasets?</title>
      <link href="https://alexgude.com/blog/where-to-host-public-datasets/" rel="alternate" type="text/html" title="Where to Host Public Datasets?" />
      <published>2021-04-26T00:00:00-07:00</published>
      <updated>2021-04-26T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/where_to_host_public_datasets</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/where-to-host-public-datasets/"><![CDATA[<p>Cleaning a dataset is tough work. I spent weeks figuring out what all the
columns of the <a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">California Statewide Integrated Traffic Records System
(SWITRS)</a> dataset meant and additional time <a href="/blog/switrs-to-sqlite/">writing scripts to parse
and fix it</a>. I wanted other people to be able to make use of the data
without going through the same hassle, so I released the scripts.</p>

<p>It wasn’t enough. People still had to request the data, download it, and then
run the scripts. Too much of a hurdle for most people. And worse, California
no longer provided some of the oldest data. It was suddenly impossible for
other people to reproduce my earlier work!</p>

<p>Luckily, I had saved all of the data. So I decided to <a href="/blog/switrs-sqlite-hosted-dataset/">host the dataset
online</a> to make it easy to start using right away.</p>

<p>But there was a problem: the dataset was so large that finding a site to host
it was not easy. In the end I choose two places: <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs"><strong>Kaggle</strong></a> and
<a href="https://zenodo.org/record/4284843"><strong>Zenodo</strong></a>. In this post I’ll share the lessons I learned, the
benefits of each site, and why I think using both is the right thing to do.</p>

<h2 id="my-requirements">My Requirements</h2>

<p>I had four requirements:</p>

<ul>
  <li>
    <p><strong>Free</strong>: The service had to be free for me, and free for the end users as
well. Requiring someone to pay for access to an open source dataset that I
had volunteered to curate is unfair and would certainly drastically reduce
the number of people making use of it.</p>
  </li>
  <li>
    <p><strong>Easy</strong>: It had to be easy for me to set up. I wanted a service where I
could get started without having to email someone for permission. I also
wanted it to be easy for the end user to get the data so that the largest
number of people could make use of it.</p>
  </li>
  <li>
    <p><strong>Discoverable</strong>: It was important to me that people could easily find the
dataset. There are dozens of sites that host large files, but they aren’t
places where people would go look for data. To help people find it I would
have to set up a web page pointing to the download, which wasn’t something I
wanted to maintain.</p>
  </li>
  <li>
    <p><strong>Permanent</strong>: Finally, there was no point in going through all this work if
it was just going to disappear tomorrow. I wanted the data to be available
for years and years.</p>
  </li>
</ul>

<p><a href="https://aws.amazon.com/opendata">AWS Open Data</a> was one option I considered, but it looked like a lot of
work to set up and it was unclear exactly how free it was. Further, getting
the data wasn’t easy if you had never worked with S3 before. Ideally, I wanted
a service that had a big button that said “Download this data!”</p>

<p>I also considered self-hosting on <a href="/blog/raspberry-pi-reboot-times/">my Raspberry Pis</a>, but quickly
dismissed it. Availability would be terrible, download speeds would be even
worse, and it would force me to perform a lot of maintenance to keep it
running.</p>

<p>In the end I settled on two services:</p>

<h2 id="kaggle">Kaggle</h2>

<p>Kaggle is a great place to host a dataset. It’s free, easy to use (with the
exception that you need an account), and is a well known place to find
datasets making it discoverable. The one downside is that Google is <a href="https://killedbygoogle.com/">infamous
for killing services</a>, so the data might not last.</p>

<p>Kaggle allows users to download the file <em>or</em> work with the data directly in
Kaggle’s hosted notebooks. The author can even set up a <a href="https://www.kaggle.com/alexgude/starter-california-traffic-collisions-from-switrs">demo
notebook</a> to demonstrate how to work with the data. Kaggle will even
help you set up a <a href="https://en.wikipedia.org/wiki/Digital_object_identifier">DOI</a> for your data. Mine is:
<a href="https://www.doi.org/10.34740/kaggle/dsv/1671261">10.34740/kaggle/dsv/1671261</a></p>

<p>Kaggle supports really deep data documentation. You can write an introduction
for each table and each column. Additionally, Kaggle will automatically
generate histograms of each column and some summary statistics.</p>

<p>Kaggle is more than just hosting; it is a community. Other people can share
their work, set up challenges using the data, and ask questions in the forum.
This community makes it easy for people to find the data and get started
working with it.</p>

<h2 id="zenodo">Zenodo</h2>

<p>Zenodo—hosted by CERN—is a much simpler and smaller service than Kaggle.
It does not have a community built up around it. It does not have attached
cloud compute. They have almost no users.<sup style="anchor-name:--fnref-usage" id="fnref:usage"><a href="#fn:usage" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>So why use Zenodo? Simple: Google <a href="https://killedbygoogle.com/">kills products left and
right</a> while CERN knows a little something about <a href="http://info.cern.ch/">keeping
websites online</a>. I trust CERN’s stewardship of the dataset. I am
far more confident that you will be able to download it from Zenodo in 10
years than from Kaggle.</p>

<p>Zenodo has great support for academic dataset usage. It allows you to use any
valid DOI, so I was able to reuse the one from Kaggle, although it will also
generate one for you if you wish. It will track citations to your dataset and
provides links to the papers. It will even let you export the citation to
<a href="https://en.wikipedia.org/wiki/BibTeX">BibTex</a> or generate a text citation on the website. Zenodo lets you
link your identity to your <a href="https://en.wikipedia.org/wiki/ORCID">Open Researcher and Contributor ID</a>.</p>

<p>Unlike Kaggle, Zenodo makes downloading easy. You do not need an account, you
just go to the page and click the download button. It also shows a <a href="https://en.wikipedia.org/wiki/MD5">MD5
hash</a> of the file so you can verify your download is exactly the same as
the file on the server.</p>

<p>Zenodo has a problem in addition to its low usage though: uploading a dataset
often fails. I originally tried to upload the uncompressed database but that
failed multiple times. After reaching out to their support (who were very
responsive and helpful), I compressed the database tried again. The smaller
file succeed where the larger one had failed.</p>

<h2 id="conclusion">Conclusion</h2>

<p>I think using both Kaggle and Zenodo is the perfect way to host a public
dataset. Kaggle has a great community and lets people quickly discover your
dataset and make use of it. The downside is the uncertain longevity and the
fact that you need an account to download the dataset. Zenodo perfectly
complements Kaggle’s weakness as its backed by CERN, an organization that
takes data hosting seriously, and makes it very easy to download the dataset.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:usage">

      <p>As of this post, my dataset has been viewed 161 times on Zenodo and
downloaded just 47 times. It has been viewed 47,800 times on Kaggle and
downloaded 3017 times. In addition, there have been 24 Kaggle notebooks
posted that make use of the data. <a href="#fnref:usage" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[When I released the SWITRS dataset, I had to find a place to host a 5 Gig dataset. Here is what I learned.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-dataset/doe_computer.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-dataset/doe_computer.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Jupyter Notebook Templates for Data Science: Plotting Time Series</title>
      <link href="https://alexgude.com/blog/data-science-timeseries-plotting-notebook-template/" rel="alternate" type="text/html" title="Jupyter Notebook Templates for Data Science: Plotting Time Series" />
      <published>2021-03-14T00:00:00-08:00</published>
      <updated>2021-03-14T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/data_science_timeseries_plotting_notebook_template</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-timeseries-plotting-notebook-template/"><![CDATA[<p>I often have data where each row describes an event. The data might describe
<a href="/blog/my-sons-language-development-comparison/#development">a word that my son spoke for the first time</a>, or <a href="/blog/switrs-bicycle-crashes-by-date/#crashes-per-week">a collision
that happened in California</a>, or <a href="/blog/2020-tour-de-france-plot/#the-race-for-yellow">the finishing place of a rider in
the Tour de France</a>. A question I always want to answer with the
data is: <em>What does the distribution of these events look like in time?</em></p>

<p>Plotting the data as a time series is the best way to answer this question,
but I never remember how to pivot the table, aggregate the events by type, and
resample to the right frequency. So I made the <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/master/notebooks/basic-time-series-plotting-template.ipynb"><strong>Time Series Plotting
Notebook</strong></a> to remember for me.</p>

<h2 id="the-time-series-plotting-notebook">The Time Series Plotting Notebook</h2>

<p>Suppose we are looking at the number of automobile collisions by make using my
<a href="/blog/switrs-sqlite-hosted-dataset/">curated SWITRS dataset</a>. We could extract one row for each
collision and the associated vehicle, which would look like this:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">ID</th>
      <th style="text-align: right">datetime</th>
      <th style="text-align: right">vehicle_make</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">2020-01-01</td>
      <td style="text-align: right">Honda</td>
    </tr>
    <tr>
      <td style="text-align: left">1</td>
      <td style="text-align: right">2020-02-01</td>
      <td style="text-align: right">Toyota</td>
    </tr>
    <tr>
      <td style="text-align: left">2</td>
      <td style="text-align: right">2020-01-01</td>
      <td style="text-align: right">Other</td>
    </tr>
    <tr>
      <td style="text-align: left">…</td>
      <td style="text-align: right">…</td>
      <td style="text-align: right">…</td>
    </tr>
  </tbody>
</table>

<p>The <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/master/notebooks/basic-time-series-plotting-template.ipynb"><strong>time series plotting notebook</strong></a> has two helpful functions
to visualize this data: <code class="language-plaintext highlighter-rouge">plot_time_series()</code> and <code class="language-plaintext highlighter-rouge">draw_left_legend()</code>.</p>

<h3 id="plot-time-series">Plot Time Series</h3>

<p>The first function, <code class="language-plaintext highlighter-rouge">plot_time_series()</code> is simple. It takes a dataframe
formatted like the above data and returns a plot showing the number of events
for each value in the categorical column. For example, to plot the number of
accidents per week by vehicle make, we would call:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nf">plot_time_series</span><span class="p">(</span>
  <span class="n">df</span><span class="p">,</span>
  <span class="n">ax</span><span class="p">,</span>
  <span class="n">date_col</span><span class="o">=</span><span class="sh">"</span><span class="s">datetime</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">category_col</span><span class="o">=</span><span class="sh">"</span><span class="s">vehicle_make</span><span class="sh">"</span><span class="p">,</span>
  <span class="n">resample_frequency</span><span class="o">=</span><span class="sh">"</span><span class="s">W</span><span class="sh">"</span><span class="p">,</span>  <span class="c1"># Resample to 'W'eeks
</span><span class="p">)</span>
</code></pre></div></div>

<p>Which would produce this plot:</p>

<p><a href="/files/jupyter-library//make_collision_in_time_first_version.svg"><img src="/files/jupyter-library//make_collision_in_time_first_version.svg" alt="A simple plot of the number of collisions by vehicle make in
California" /></a></p>

<p>The function accepts a few optional parameters:</p>

<ul>
  <li>
    <p><code class="language-plaintext highlighter-rouge">resample_frequency</code>: controls the timescale over which the data is
aggregated.</p>
  </li>
  <li>
    <p><code class="language-plaintext highlighter-rouge">aggfunc</code> which controls how the data is aggregated.</p>
  </li>
  <li>
    <p><code class="language-plaintext highlighter-rouge">linewidth</code> which can be used to make the lines larger if there are only a
few of them, or thinner if there is lots of data.</p>
  </li>
</ul>

<h3 id="simple-legend">Simple Legend</h3>

<p>Simple legends are great. They convey their information effectively because
the superfluous noise has been removed. My <a href="/blog/data-science-plotting-notebook-template/#draw-legends">basic plotting
notebook</a> has a function to remove all the extra information
from the legend box leaving only the color and the label. This time I have
taken it a step further: I wrote a function to get rid of the box and label
each line.</p>

<p>The function <code class="language-plaintext highlighter-rouge">draw_left_legend()</code> will draw labels on the end of each line,
like so:</p>

<p><a href="/files/jupyter-library//make_collision_in_time_second_version.svg"><img src="/files/jupyter-library//make_collision_in_time_second_version.svg" alt="A simple plot of the number of collisions by vehicle make in California
with left legend" /></a></p>

<p>I’ve used this legend when <a href="/blog/2019-tour-de-france-plot/#the-race-for-yellow"><em>Plotting the winners of the 2019 Tour de
France</em></a> as well as the <a href="/blog/2020-tour-de-france-plot/#the-race-for-yellow"><em>2020 Tour de France</em></a>.</p>

<h2 id="putting-it-together">Putting It Together</h2>

<p>The <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/master/notebooks/basic-time-series-plotting-template.ipynb">time series plotting notebook</a> enables you to quickly plot
your data in time with only a few lines of code. Here is the final version of
the plot:</p>

<p><a href="/files/jupyter-library//make_collision_in_time.svg"><img src="/files/jupyter-library//make_collision_in_time.svg" alt="An example plot from the notebook library" /></a></p>

<p>Which was produced by this short code snippet:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="n">seaborn</span> <span class="k">as</span> <span class="n">sns</span>

<span class="n">fig</span><span class="p">,</span> <span class="n">ax</span> <span class="o">=</span> <span class="nf">setup_plot</span><span class="p">(</span><span class="n">title</span><span class="o">=</span><span class="sh">"</span><span class="s">Collisions by Make</span><span class="sh">"</span><span class="p">)</span>

<span class="n">pivot</span> <span class="o">=</span> <span class="nf">plot_time_series</span><span class="p">(</span><span class="n">df</span><span class="p">,</span> <span class="n">ax</span><span class="p">,</span> <span class="n">date_col</span><span class="o">=</span><span class="n">DATE_COL</span><span class="p">,</span> <span class="n">category_col</span><span class="o">=</span><span class="sh">"</span><span class="s">vehicle_make</span><span class="sh">"</span><span class="p">,</span> <span class="n">resample_frequency</span><span class="o">=</span><span class="sh">"</span><span class="s">W</span><span class="sh">"</span><span class="p">)</span>

<span class="c1"># Move labels slightly to avoid overlap
</span><span class="n">nudges</span> <span class="o">=</span> <span class="p">{</span><span class="sh">"</span><span class="s">Toyota</span><span class="sh">"</span><span class="p">:</span> <span class="mi">15</span><span class="p">,</span> <span class="sh">"</span><span class="s">Honda</span><span class="sh">"</span><span class="p">:</span> <span class="o">-</span><span class="mi">8</span><span class="p">}</span>
<span class="nf">draw_left_legend</span><span class="p">(</span><span class="n">ax</span><span class="p">,</span> <span class="n">nudges</span><span class="o">=</span><span class="n">nudges</span><span class="p">,</span> <span class="n">fontsize</span><span class="o">=</span><span class="mi">25</span><span class="p">)</span>

<span class="n">sns</span><span class="p">.</span><span class="nf">despine</span><span class="p">(</span><span class="n">trim</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>

<span class="nf">save_plot</span><span class="p">(</span><span class="n">fig</span><span class="p">,</span> <span class="sh">"</span><span class="s">/tmp/make_collision_in_time.svg</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>I hope the notebook template library is useful to you! Let me know on
<a href="https://twitter.com/alex_gude/">Twitter</a> or <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/issues">Github</a> if it is. Your feedback helps make the
project better for everyone!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
          <category term="data-science" />
        
          <category term="jupyter" />
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[Jumpstart your time series visualizations with this Jupyter plotting notebook!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/jupyter-library/jupiter_in_the_rearview_mirror.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/jupyter-library/jupiter_in_the_rearview_mirror.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Making Custom Markdown for Github Pages</title>
      <link href="https://alexgude.com/blog/custom-markdown-for-github-pages/" rel="alternate" type="text/html" title="Making Custom Markdown for Github Pages" />
      <published>2021-02-08T00:00:00-08:00</published>
      <updated>2021-02-08T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/custom_markdown_for_github_pages</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/custom-markdown-for-github-pages/"><![CDATA[<p><a href="https://en.wikipedia.org/wiki/Markdown">Markdown</a> is a great way to write; simple enough to be read as text, but with
the ability to fall back to full HTML if required. I have been using it for my
own notes for a decade and this site is written in Markdown using <a href="https://en.wikipedia.org/wiki/Jekyll_(software)">Jekyll</a>.</p>

<p>Markdown provides a lot of syntax to simplify HTML, like <code class="language-plaintext highlighter-rouge">**BOLD**</code> to create
<code class="language-plaintext highlighter-rouge">&lt;strong&gt;BOLD&lt;/strong&gt;</code> text, or <code class="language-plaintext highlighter-rouge">&gt; Quote</code> to create
<code class="language-plaintext highlighter-rouge">&lt;blockquote&gt;Quote&lt;/blockquote&gt;</code>. Recently, for my <a href="https://github.com/MiniFate/MiniFate">MiniFate</a> project, I
wanted to add a few custom <code class="language-plaintext highlighter-rouge">&lt;span&gt;</code> elements to highlight specific pieces of
the text. I could have fallen back to writing it out in HTML each time, but
doing so felt clunky when compared to how smooth writing Markdown normally is.</p>

<p>I created a way to define my own syntax based on <a href="https://bro.doktorbro.net/">Anatol Broder’s</a>
<a href="https://jch.penibelst.de/">Compress</a> and <a href="https://sylvaindurand.org/">Sylvain Durand’s</a> post on <a href="https://sylvaindurand.org/improving-typography-on-jekyll/"><em>Improving typography on
Jekyll</em></a>. It uses <a href="https://shopify.github.io/liquid/">Liquid</a> to rewrite the web page <strong>after</strong> it has been
compiled, giving you complete control of formatting and allowing you to define
custom Markdown syntax. Since it uses only the default tools built in to
Jekyll, it works natively on <a href="https://pages.github.com/">Github Pages</a>! Here is how it works:</p>

<h2 id="layouts">Layouts</h2>

<p>The <a href="https://jekyllrb.com/tutorials/orderofinterpretation/">order of interpretation</a> to build a page in Jekyll is:</p>

<ol>
  <li>
    <p>Substitute site and page variables.</p>
  </li>
  <li>
    <p>Run Liquid functions.</p>
  </li>
  <li>
    <p>Compile Markdown to HTML.</p>
  </li>
  <li>
    <p>Push the compiled HTML into its layout template, if there is one.</p>
  </li>
  <li>
    <p>Write to file.</p>
  </li>
</ol>

<p>The problem is that Liquid runs before compiling, so we can’t use Liquid code
embedded on a page to modify the final HTML. But there is a workaround: the
fully compiled HTML is pushed to a layout (if one is specified) and that
layout restarts the page build order! This means we can modify a page’s HTML
using Liquid written in the page’s layout template!</p>

<p>To define our custom syntax then, we just need to write a simple layout and
place it in <code class="language-plaintext highlighter-rouge">_layouts/substitute.html</code>:</p>

<div class="language-liquid highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">{%</span><span class="w"> </span><span class="nt">comment</span><span class="w"> </span><span class="cp">%}</span><span class="c">
&lt;!-- This is the code block to define custom syntax --&gt;
</span><span class="cp">{%</span><span class="w"> </span><span class="nt">endcomment</span><span class="w"> </span><span class="cp">%}</span>
<span class="cp">{%</span><span class="w"> </span><span class="nt">assign</span><span class="w"> </span><span class="nv">output</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">content</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'-!'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;u&gt;'</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'!-'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;/u&gt;'</span><span class="w">
</span><span class="cp">%}</span>

<span class="cp">{{</span><span class="nv">output</span><span class="cp">}}</span>
</code></pre></div></div>

<p>Here <code class="language-plaintext highlighter-rouge">content</code> is the special variable that contains the compiled HTML from
the page that is using the template.</p>

<p>We then change our primary layout (probably <code class="language-plaintext highlighter-rouge">_layouts/default.html</code>) to
inherit from <code class="language-plaintext highlighter-rouge">substitute</code>:</p>

<div class="language-html highlighter-rouge"><div class="highlight"><pre class="highlight"><code>---
layout: substitute
---

<span class="cp">&lt;!DOCTYPE html&gt;</span>
<span class="nt">&lt;html</span> <span class="na">lang=</span><span class="s">"en"</span><span class="nt">&gt;</span>
  <span class="nt">&lt;body&gt;</span>
    <span class="nt">&lt;main&gt;</span>{{ content }}<span class="nt">&lt;/main&gt;</span>
  <span class="nt">&lt;/body&gt;</span>
<span class="nt">&lt;/html&gt;</span>
</code></pre></div></div>

<p>And that’s it! All the customization is controlled by changing the Liquid code
in <code class="language-plaintext highlighter-rouge">substitute.html</code>. Below are some examples.</p>

<h3 id="defining-custom-markup">Defining Custom Markup</h3>

<p>Markdown has no syntax for <u>Underline</u>, but we can define some like this:</p>

<div class="language-liquid highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">{%</span><span class="w"> </span><span class="nt">assign</span><span class="w"> </span><span class="nv">output</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">content</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'-!'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;u&gt;'</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'!-'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;/u&gt;'</span><span class="w">
</span><span class="cp">%}</span>
</code></pre></div></div>

<p>Now <code class="language-plaintext highlighter-rouge">-!Underline!-</code> compiles to <code class="language-plaintext highlighter-rouge">&lt;u&gt;Underline&lt;/u&gt;</code>.</p>

<p>But we can go further: we can define anything we’d like in the substitution,
for example a <code class="language-plaintext highlighter-rouge">&lt;span&gt;</code>:</p>

<div class="language-liquid highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">{%</span><span class="w"> </span><span class="nt">assign</span><span class="w"> </span><span class="nv">output</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">content</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'-!'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;span class="book-title"&gt;'</span><span class="w">
    </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'!-'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;/span&gt;'</span><span class="w">
</span><span class="cp">%}</span>
</code></pre></div></div>

<p>Which can be fully customized with CSS.</p>

<p>This method has two limitations:</p>

<ul>
  <li>
    <p>We have to use characters that the Markdown parser won’t
interpret.<sup style="anchor-name:--fnref-reserved" id="fnref:reserved"><a href="#fn:reserved" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>
  </li>
  <li>
    <p>We must define unique opening and closing syntax to match the opening and
closing HTML elements.</p>
  </li>
</ul>

<p>We can avoid these constraints by overriding standard Markdown syntax.</p>

<h3 id="overriding-markdown-syntax">Overriding Markdown Syntax</h3>

<p>I never use <code class="language-plaintext highlighter-rouge">~~Strike~~</code> in my writing, which inserts <code class="language-plaintext highlighter-rouge">&lt;del&gt;Strike&lt;/del&gt;</code> to
denote text that has been removed. We can override it to insert
<code class="language-plaintext highlighter-rouge">&lt;u&gt;Underline&lt;/u&gt;</code> instead as follows:</p>

<div class="language-liquid highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">{%</span><span class="w"> </span><span class="nt">assign</span><span class="w"> </span><span class="nv">output</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">content</span><span class="w">
  </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'&lt;del&gt;'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;u&gt;'</span><span class="w">
  </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'&lt;/del&gt;'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;/u&gt;'</span><span class="w">
</span><span class="cp">%}</span>
</code></pre></div></div>

<p>Notice that I didn’t replace <code class="language-plaintext highlighter-rouge">~~</code>, I replaced <code class="language-plaintext highlighter-rouge">&lt;del&gt;</code>. This is because the
template Liquid substitutes <em>after</em> the Markdown is compiled to HTML, and so
<code class="language-plaintext highlighter-rouge">~~</code> has already been removed from the page.</p>

<h3 id="replacement">Replacement</h3>

<p>Of course, we can replace <strong>anything</strong> using this method, not just custom
Markdown syntax or HTML elements. We can define a macro that is replaced by an
image or table. We can even reshape the page, for example, adding an <code class="language-plaintext highlighter-rouge">&lt;hr&gt;</code>
above the footnotes automatically like this:</p>

<div class="language-liquid highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">{%</span><span class="w"> </span><span class="nt">assign</span><span class="w"> </span><span class="nv">output</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">content</span><span class="w">
  </span><span class="p">|</span><span class="w"> </span><span class="nf">replace</span><span class="p">:</span><span class="w"> </span><span class="s1">'&lt;div class="footnotes" role="doc-endnotes"&gt;'</span><span class="p">,</span><span class="w"> </span><span class="s1">'&lt;hr&gt;&lt;div class="footnotes" role="doc-endnotes"&gt;'</span><span class="w">
</span><span class="cp">%}</span>
</code></pre></div></div>

<h2 id="conclusion">Conclusion</h2>

<p>Using layouts gives us the full power of <a href="https://shopify.github.io/liquid/">Liquid</a> to update our web pages
after the HTML is compiled, and it works natively on Github pages! I hope you
use this to build awesome web pages and if you do let me know on BlueSky: <a href="https://bsky.app/profile/alexgude.com">@alexgude.com</a></p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:reserved">

      <p>This means you can’t use <code class="language-plaintext highlighter-rouge">_</code>, <code class="language-plaintext highlighter-rouge">*</code>, <code class="language-plaintext highlighter-rouge">`</code>, and <code class="language-plaintext highlighter-rouge">~~</code>. Some syntax
characters like <code class="language-plaintext highlighter-rouge">#</code>, <code class="language-plaintext highlighter-rouge">&gt;</code>, and <code class="language-plaintext highlighter-rouge">-</code> are OK as long as they aren’t used at
the start of a line. <a href="#fnref:reserved" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[I love Markdown, I take all my notes in it and write my blog in it. But sometimes you want to create new syntax; read on to find out how!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/jekyll/stansaab_censor_and_operator.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/jekyll/stansaab_censor_and_operator.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Data Science Interview Practice: Machine Learning Case Study</title>
      <link href="https://alexgude.com/blog/data-science-interview-prep-case-study/" rel="alternate" type="text/html" title="Data Science Interview Practice: Machine Learning Case Study" />
      <published>2021-01-18T00:00:00-08:00</published>
      <updated>2021-01-18T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/data_science_interview_prep_case_study</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-interview-prep-case-study/"><![CDATA[<p>A common interview type for data scientists and machine learning engineers is
the machine learning case study. In it, the interviewer will ask a question
about how the candidate would build a certain model. These questions can be
challenging for new data scientists because the interview is open-ended and
new data scientists often lack practical experience building and shipping
product-quality models.</p>

<p>I have a lot of practice with these types of interviews as a result of my time
at <a href="/blog/should-i-go-to-insight/">Insight</a>, my many experiences <a href="/blog/interviewing-for-data-science-positions-in-2020/">interviewing for
jobs</a>, and my role in designing and implementing Intuit’s data
science interview. Similar to my last article where I <a href="/blog/data-science-interview-prep-data-manipulation/">put together an example
data manipulation interview practice problem</a>, this time I will
walk through a practice case study and how I would work through it.</p>

<h2 id="my-approach">My Approach</h2>

<p>Case study interviews are just conversations. This can make them tougher than
they need to be for junior data scientists because they lack the obvious
structure of a coding interview or <a href="/blog/data-science-interview-prep-data-manipulation/">data manipulation interview</a>. I
find it’s helpful to impose my own structure on the conversation by
approaching it in this order:</p>

<ol>
  <li>
    <p><strong>Problem</strong>: Dive in with the interviewer and explore what the problem is.
Look for edge cases or simple and high-impact parts of the problem that you
might be able to close out quickly.</p>
  </li>
  <li>
    <p><strong>Metrics</strong>: Once you have determined the scope and parameters of the
problem you’re trying to solve, figure out how you will measure success.
Focus on what is important to the business and not just what is easy to
measure.</p>
  </li>
  <li>
    <p><strong>Data</strong>: Figure out what data is available to solve the problem. The
interviewer might give you a couple of examples, but ask about additional
information sources. If you know of some public data that might be useful,
bring it up here too.</p>
  </li>
  <li>
    <p><strong>Labels and Features</strong>: Using the data sources you discussed, what
features would you build? If you are attacking a supervised classification
problem, how would you generate labels? How would you see if they were
useful?</p>
  </li>
  <li>
    <p><strong>Model</strong>: Now that you have a metric, data, features, and labels, what
model is a good fit? Why? How would you train it? What do you need to watch
out for?</p>
  </li>
  <li>
    <p><strong>Validation</strong>: How would you make sure your model works offline? What data
would you hold out to test your model works as expected? What metrics would
you measure?</p>
  </li>
  <li>
    <p><strong>Deployment and Monitoring</strong>: Having developed a model you are comfortable
with, how would you deploy it? Does it need to be real-time or is it
sufficient to batch inputs and periodically run the model? How would you
check performance in production? How would you monitor for model drift
where its performance changes over time?</p>
  </li>
</ol>

<h2 id="case-study">Case Study</h2>

<p>Here is the prompt:</p>

<blockquote>
  <p>At Twitter, bad actors occasionally use automated accounts, known as “bots”,
to abuse our platform. How would you build a system to help detect bot
accounts?</p>
</blockquote>

<h3 id="problem">Problem</h3>

<p>At the start of the interview I try to fully explore the bounds of the
problem, which is often open ended. My goal with this part of the interview is
to:</p>

<ul>
  <li>
    <p>Understand the problem and all the edge cases.</p>
  </li>
  <li>
    <p>Come to an agreement with the interviewer on the scope—narrower is
better!—of the problem to solve.</p>
  </li>
  <li>
    <p>Demonstrate any knowledge I have on the subject, especially from researching
the company previously.</p>
  </li>
</ul>

<p>Our Twitter bot prompt has a lot of angles from which we could attack. I know
Twitter has dozens of types of bots, ranging from my <a href="/blog/raspberry-pi-reboot-times/">harmless Raspberry Pi
bots</a>, to <a href="https://en.wikipedia.org/wiki/Russian_web_brigades">“Russian Bots” trying to influence elections</a>,
to <a href="https://en.wikipedia.org/wiki/Spambot">bots spreading spam</a>. I would pick one problem to focus on using
my best guess as to business impact. In this case spam bots are likely a
problem that causes measurable harm (drives users away, drives advertisers
away). Russian bots are probably a bigger issue in terms of public perception,
but that’s much harder to measure.</p>

<p>After deciding on the scope, I would ask more about the systems they currently
have to deal with it. Likely Twitter has an ops team to help identify spam and
block accounts and they may even have a rules based system. Those systems will
be a good source of data about the bad actors and they likely also have
metrics they track for this problem.</p>

<h3 id="metric">Metric</h3>

<p>Having agreed on what part of the problem to focus on, we now turn to how we
are going to measure our impact. There is no point shipping a model if you
can’t measure how it’s affecting the business.</p>

<p>Metrics and model use go hand-in-hand, so first we have to agree on what the
model will be used for. For spam we could use the model to just mark suspected
accounts for human review and tracking, or we could outright block accounts
based on the model result. If we pick the human review option, it’s probably
more important to get all the bots even if some good customers are affected.
If we go with immediate action, it is likely more important to only ban truly
bad accounts. I covered thinking about metrics like this in detail in another
post, <a href="/blog/machine-learning-metrics-interview/"><em>What Machine Learning Metric to Use</em></a>. Take a look!</p>

<p>I would argue the automatic blocking model will have higher impact because it
frees our ops people to focus on other bad behavior. We want two sets of
metrics: <strong>offline</strong> for when we are training and <strong>online</strong> for when the
model is deployed.</p>

<p>Our offline metric will be <strong>precision</strong> because, based on the argument above,
we want to be really sure we’re only banning bad accounts.</p>

<p>Our online metrics are more business focused:</p>

<ul>
  <li>
    <p><strong>Ops time saved</strong>: Ops is currently spending some amount of time reviewing
spam; how much can we cut that down?</p>
  </li>
  <li>
    <p><strong>Spam fraction</strong>: What percent of Tweets are spam? Can we reduce this?</p>
  </li>
</ul>

<p>It is often useful to normalize metrics, like the spam fraction metric, so
they don’t go up or down just because we have more customers!</p>

<h3 id="data">Data</h3>

<p>Now that we know what we’re doing and how to measure its success, it’s time to
figure out what data we can use. Just based on how a company operates, you can
make a really good guess as to the data they have. For Twitter we know they
have to track Tweets, accounts, and logins, so they must have databases with
that information. Here are what I think they contain:</p>

<ul>
  <li>
    <p><strong>Tweets database</strong>: Sending account, mentioned accounts, parent Tweet,
Tweet text.</p>
  </li>
  <li>
    <p><strong>Interactions database</strong>: Account, Tweet, action (retweet, favorite, etc.).</p>
  </li>
  <li>
    <p><strong>Accounts database</strong>: Account name, handle, creation date, creation
device, creation IP address.</p>
  </li>
  <li>
    <p><strong>Following database</strong>: Account, followed account.</p>
  </li>
  <li>
    <p><strong>Login database</strong>: Account, date, login device, login IP address, success
or fail reason.</p>
  </li>
  <li>
    <p><strong>Ops database</strong>: Account, restriction, human reasoning.</p>
  </li>
</ul>

<p>And a lot more. From these we can find out a lot about an account and the
Tweets they send, who they send to, who those people react to, and possibly
how login events tie different accounts together.</p>

<h3 id="labels-and-features">Labels and Features</h3>

<p>Having figured out what data is available, it’s time to process it. Because
I’m treating this as a classification problem, I’ll need <strong>labels</strong> to tell me
the ground truth for accounts, and I’ll need <strong>features</strong> which describe the
behavior of the accounts.</p>

<h3 id="labels">Labels</h3>

<p>Since there is an ops team handling spam, I have historical examples of bad
behavior which I can use as positive labels.<sup style="anchor-name:--fnref-positive_labels" id="fnref:positive_labels"><a href="#fn:positive_labels" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> If there aren’t
enough I can use tricks to try to expand my labels, for example looking at IP
address or devices that are associated with spammers and labeling other
accounts with the same login characteristics.</p>

<p>Negative labels are harder to come by. I know Twitter has verified users who
are unlikely to be spam bots, so I can use them. But verified users are
certainly very different from “normal” good users because they have far more
followers.</p>

<p>It is a safe bet that there are far more good users than spam bots, so
randomly selecting accounts can be used to build a negative label set.</p>

<h3 id="features">Features</h3>

<p>To build features, it helps to think about what sort of behavior a spam bot
might exhibit, and then try to codify that behavior into features. For
example:</p>

<ul>
  <li>
    <p><strong>Bots can’t write truly unique messages</strong>; they must use a template or
language generator. This should lead to similar messages, so looking at how
repetitive an account’s Tweets are is a good feature.</p>
  </li>
  <li>
    <p><strong>Bots are used because they scale.</strong> They can run all the time and send
messages to hundreds or thousands (or millions) or users. Number of unique
Tweet recipients and number of minutes per day with a Tweet sent are likely
good features.</p>
  </li>
  <li>
    <p><strong>Bots have a controller.</strong> Someone is benefiting from the spam, and they
have to control their bots. Features around logins might help here like
number of accounts seen from this IP address or device, similarity of login
time, etc.</p>
  </li>
</ul>

<h3 id="model-selection">Model Selection</h3>

<p>I try to start with the simplest model that will work when starting a new
project. Since this is a supervised classification problem and I have written
some simple features, logistic regression or a forest are good candidates. I
would likely go with a forest because they tend to “just work” and are a
little less sensitive to feature processing.<sup style="anchor-name:--fnref-processing" id="fnref:processing"><a href="#fn:processing" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>

<p>Deep learning is not something I would use here. It’s great for image, video,
audio, or NLP, but for a problem where you have a set of labels and a set of
features that you believe to be predictive it is generally overkill.</p>

<p>One thing to consider when training is that the dataset is probably going to
be wildly imbalanced. I would start by down-sampling (since we likely have
millions of events), but would be ready to discuss other methods and trade
offs.</p>

<h3 id="validation">Validation</h3>

<p>Validation is not too difficult at this point. We focus on the offline metric
we decided on above: precision. We don’t have to worry much about leaking data
between our holdout sets if we split at the account level, although if we
include bots from the same <a href="https://en.wikipedia.org/wiki/Botnet">botnet</a> into our different sets there will
be a little data leakage. I would start with a simple validation/training/test
split with fixed fractions of the dataset.</p>

<h3 id="deployment">Deployment</h3>

<p>Since we want to classify an entire account and not a specific tweet, we don’t
need to run the model in real-time when Tweets are posted. Instead we can run
batches and can decide on the time between runs by looking at something like
the characteristic time a spam bot takes to send out Tweets. We can add rate
limiting to Tweet sending as well to slow the spam bots and give us more time
to decide without impacting normal users.</p>

<p>For deployment, I would start in <strong>shadow mode</strong>, which I <a href="/blog/machine-learning-deployment-shadow-mode/">discussed in detail
in another post</a>. This would allow us to see how the model
performs on real data without the risk of blocking good accounts. I would
track its performance using our online metrics: spam fraction and ops time
saved. I would compute these metrics twice, once using the assumption that the
model blocks flagged accounts, and once assuming that it does not block
flagged accounts, and then compare the two outcomes. If the comparison is
favorable, the model should be promoted to action mode.</p>

<h2 id="let-me-know">Let Me Know!</h2>

<p>I hope this exercise has been helpful! Please reach out and let me know on
BlueSky at <a href="https://bsky.app/profile/alexgude.com">@alexgude.com</a> if you have any comments or
improvements!</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:positive_labels">

      <p>In this case a <em>positive label</em> means the account is a spam bot, and a
<em>negative label</em> means they are not. <a href="#fnref:positive_labels" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:processing">

      <p>If you use <a href="https://en.wikipedia.org/wiki/Regularization_(mathematics)">regularization</a> with logistic regression (and you should) you
need to scale your features. Random forests do not require this. <a href="#fnref:processing" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="interviewing" />
        
          <category term="interview-prep" />
        
      

      

      
      
        <summary type="html"><![CDATA[A common interview type for data scientists and machine learning engineers is the ML case study. Read on for an example of how I solve them!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/interview-prep/henry_reid_at_his_desk_nasa.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/interview-prep/henry_reid_at_his_desk_nasa.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Data Science Interview Practice: Data Manipulation</title>
      <link href="https://alexgude.com/blog/data-science-interview-prep-data-manipulation/" rel="alternate" type="text/html" title="Data Science Interview Practice: Data Manipulation" />
      <published>2020-12-14T00:00:00-08:00</published>
      <updated>2020-12-14T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/data_science_interview_prep_data_manipulation</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-interview-prep-data-manipulation/"><![CDATA[<p>I often get asked by newly-minted PhDs trying to get their first data job:</p>

<blockquote>
  <p>How can I prepare for dataset-based interviews? Do you have any examples of
datasets to practice with?</p>
</blockquote>

<p>I never had a good answer. I would tell them about how the interviews worked,
but I wished I had something to share that they could get their hands on.</p>

<p>As of today, that’s changing. In this post I put together a series of practice
questions like the kind you might see (or be expected to come up with) in a
hands-on data interview using the <a href="/blog/switrs-sqlite-hosted-dataset/">curated and hosted dataset of California
Traffic accidents</a>. The dataset is available for download from
both <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs">Kaggle</a> and <a href="https://zenodo.org/record/4284843">Zenodo</a>, and I even have an <a href="https://www.kaggle.com/alexgude/starter-california-traffic-collisions-from-switrs">example
notebook</a> demonstrating how to work with the data entirely
online within Kaggle.</p>

<h2 id="interview-format">Interview Format</h2>

<p>As I mentioned in <a href="/blog/interviewing-for-data-science-positions-in-2020/">my post about my most recent interview
experience</a>, data science and machine learning interviews have
become more practical, covering tasks that show up in the day-to-day work of a
data scientist instead of hard but irrelevant problems. One common interview
type involves working with a dataset, answering some simple questions about
it, and then building some simple features.</p>

<p>Generally these interviews use Python and <a href="https://en.wikipedia.org/wiki/Pandas_(software)">Pandas</a> or pure SQL.
Sometimes the interviewer has a set of questions for you to answer and
sometimes they want you to come up with your own.</p>

<p>To help people prepare, I have created a set of questions similar to what you
would get in a real interview. For the exercise you will be using the SWITRS
dataset. I have included a notebook to get you started in <a href="/files/interview-prep/Interview%20Prep%20SQL.ipynb"><strong>SQL</strong></a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/interview-prep/Interview%20Prep%20SQL.ipynb">rendered on Github</a>) or <a href="/files/interview-prep/Interview%20Prep%20Python.ipynb"><strong>Pandas</strong></a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/interview-prep/Interview%20Prep%20Python.ipynb">rendered on Github</a>). The solution notebooks can be
found at the very end and I have included answers in this post, click the
arrow to show them.</p>

<p>Good luck, and if you have any questions or suggestions please reach out to me
on BlueSky: <a href="https://bsky.app/profile/alexgude.com">@alexgude.com</a></p>

<h2 id="questions">Questions</h2>

<h3 id="how-many-collisions-are-there-in-the-dataset">How many collisions are there in the dataset?</h3>

<details>

  <summary>
    <p>A good first thing to check is “How much data am I dealing with?”</p>
  </summary>

  <p>Each row in the collisions database represents one collision, so the solution
is nice and short:</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="k">COUNT</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">AS</span> <span class="n">collision_count</span>
<span class="k">FROM</span> <span class="n">collisions</span>
</code></pre></div>  </div>

  <p>Which returns:</p>

  <div class="low-width-table" style="max-width: 20%">

    <table>
      <thead>
        <tr>
          <th style="text-align: right">collision_count</th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td style="text-align: right">9,172,565</td>
        </tr>
      </tbody>
    </table>

  </div>
</details>

<h3 id="what-percent-of-collisions-involve-males-aged-1625">What percent of collisions involve males aged 16–25?</h3>

<details>

  <summary>
    <p>Young men are famously unsafe drivers so let’s look at how many collisions
they’re involved in.</p>
  </summary>

  <p>The age and gender of the drivers are in the parties table so the query does a
simple filter on those entries. The tricky part comes from needing to
calculate the ratio as this requires us to get the total number of collisions.
We could hard-code the number, but I prefer calculating it as part of the
query. There isn’t a super elegant way to do it in SQLite, but a sub-query
works fine. We also have to cast to a float to avoid integer division.</p>

  <p>There are a lot of <code class="language-plaintext highlighter-rouge">NULL</code> values for age and sex. I assume they are
uncorrelated to age and sex which allows me to remove them. If we were worried
about this assumption, we could leave them in and treat the answer as a lower
bound.</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span>
    <span class="k">COUNT</span><span class="p">(</span><span class="k">DISTINCT</span> <span class="n">case_id</span><span class="p">)</span>
    <span class="o">/</span> <span class="p">(</span>
        <span class="k">SELECT</span> <span class="k">CAST</span><span class="p">(</span><span class="k">COUNT</span><span class="p">(</span><span class="k">DISTINCT</span> <span class="n">case_id</span><span class="p">)</span> <span class="k">AS</span> <span class="nb">FLOAT</span><span class="p">)</span>
        <span class="k">FROM</span> <span class="n">parties</span>
        <span class="k">WHERE</span> <span class="n">party_age</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
        <span class="k">AND</span> <span class="n">party_sex</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
        <span class="p">)</span>
    <span class="k">AS</span> <span class="n">percentage</span>
<span class="k">FROM</span> <span class="n">parties</span>
<span class="k">WHERE</span> <span class="n">party_sex</span> <span class="o">=</span> <span class="s1">'male'</span>
<span class="k">AND</span> <span class="n">party_age</span> <span class="k">BETWEEN</span> <span class="mi">16</span> <span class="k">AND</span> <span class="mi">25</span>
</code></pre></div>  </div>

  <p>The result is:</p>

  <div class="low-width-table" style="max-width: 20%">

    <table>
      <thead>
        <tr>
          <th style="text-align: right">percentage</th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td style="text-align: right">0.258</td>
        </tr>
      </tbody>
    </table>

  </div>

</details>

<h3 id="how-many-solo-motorcycle-crashes-are-there-per-year">How many solo motorcycle crashes are there per year?</h3>

<details>

  <summary>
    <p>A “<em>solo</em>” crash is one where the driver runs off the road or hits a
stationary object. How many solo motorcycle crashes were there each year? Why
does 2020 seem to (relatively) have so few?</p>
  </summary>

  <p>To select the right rows we filter with <code class="language-plaintext highlighter-rouge">WHERE</code> and to get the count per year
we need to use a <code class="language-plaintext highlighter-rouge">GROUP BY</code>. SQLite does not have a <code class="language-plaintext highlighter-rouge">YEAR()</code> function, so we
have to use <code class="language-plaintext highlighter-rouge">strftime</code> instead. In a real interview, you can normally just
assume that the function you need will exist without getting into the
specifics of the SQL dialect.</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span>
  <span class="n">STRFTIME</span><span class="p">(</span><span class="s1">'%Y'</span><span class="p">,</span> <span class="n">collision_date</span><span class="p">)</span> <span class="k">AS</span> <span class="n">collision_year</span><span class="p">,</span>
  <span class="k">COUNT</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">AS</span> <span class="n">collision_count</span>
<span class="k">FROM</span> <span class="n">collisions</span>
<span class="k">WHERE</span> <span class="n">motorcycle_collision</span> <span class="o">=</span> <span class="k">True</span>
  <span class="k">AND</span> <span class="n">party_count</span> <span class="o">=</span> <span class="mi">1</span>
<span class="k">GROUP</span> <span class="k">BY</span> <span class="n">collision_year</span>
<span class="k">ORDER</span> <span class="k">BY</span> <span class="n">collision_year</span>
</code></pre></div>  </div>

  <p>This gives us:</p>

  <div class="low-width-table" style="max-width: 20%">

    <table>
      <thead>
        <tr>
          <th style="text-align: right">collision_year</th>
          <th style="text-align: right">collision_count</th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td style="text-align: right">2001</td>
          <td style="text-align: right">3258</td>
        </tr>
        <tr>
          <td style="text-align: right">2002</td>
          <td style="text-align: right">3393</td>
        </tr>
        <tr>
          <td style="text-align: right">2003</td>
          <td style="text-align: right">3822</td>
        </tr>
        <tr>
          <td style="text-align: right">2004</td>
          <td style="text-align: right">3955</td>
        </tr>
        <tr>
          <td style="text-align: right">2005</td>
          <td style="text-align: right">3755</td>
        </tr>
        <tr>
          <td style="text-align: right">2006</td>
          <td style="text-align: right">3967</td>
        </tr>
        <tr>
          <td style="text-align: right">2007</td>
          <td style="text-align: right">4513</td>
        </tr>
        <tr>
          <td style="text-align: right">2008</td>
          <td style="text-align: right">4948</td>
        </tr>
        <tr>
          <td style="text-align: right">2009</td>
          <td style="text-align: right">4266</td>
        </tr>
        <tr>
          <td style="text-align: right">2010</td>
          <td style="text-align: right">3902</td>
        </tr>
        <tr>
          <td style="text-align: right">2011</td>
          <td style="text-align: right">4054</td>
        </tr>
        <tr>
          <td style="text-align: right">2012</td>
          <td style="text-align: right">4143</td>
        </tr>
        <tr>
          <td style="text-align: right">2013</td>
          <td style="text-align: right">4209</td>
        </tr>
        <tr>
          <td style="text-align: right">2014</td>
          <td style="text-align: right">4267</td>
        </tr>
        <tr>
          <td style="text-align: right">2015</td>
          <td style="text-align: right">4415</td>
        </tr>
        <tr>
          <td style="text-align: right">2016</td>
          <td style="text-align: right">4471</td>
        </tr>
        <tr>
          <td style="text-align: right">2017</td>
          <td style="text-align: right">4373</td>
        </tr>
        <tr>
          <td style="text-align: right">2018</td>
          <td style="text-align: right">4240</td>
        </tr>
        <tr>
          <td style="text-align: right">2019</td>
          <td style="text-align: right">3772</td>
        </tr>
        <tr>
          <td style="text-align: right">2020</td>
          <td style="text-align: right">2984</td>
        </tr>
      </tbody>
    </table>

  </div>

  <p>The count is low in 2020 primarily because the data doesn’t cover the whole
year. It is also low due to the COVID pandemic keeping people off the streets,
at least initially. To differentiate these two causes we could compare month
by month to last year.</p>

</details>

<h3 id="what-make-of-vehicle-has-the-largest-fraction-of-accidents-on-the-weekend-during-the-work-week">What make of vehicle has the largest fraction of accidents on the weekend? During the work week?</h3>

<details>

  <summary>
    <p>Weekdays are generally commute and work-related traffic, while weekends
involves recreational travel. Do we see different vehicles involved in
collisions on these days?</p>

    <p>Only consider vehicle makes with at least 10,000 collisions, in order to focus
only on common vehicles where the difference between weekend and weekday usage
will be significant.</p>

  </summary>

  <p>This query is tricky. We need to aggregate collisions by vehicle make, which
means we need the parties table. We also care about when the crash happened,
which means we need the collisions table. So we need to join these two tables
together.</p>

  <p>In an interview setting, I would write two simpler queries: one that gets the
highest weekend fraction and one that gets the highest weekday fraction with a
lot of copy and pasted code. This is a lot easier to work out. Here is an
example of one of those queries:</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span>
  <span class="n">p</span><span class="p">.</span><span class="n">vehicle_make</span> <span class="k">AS</span> <span class="n">make</span><span class="p">,</span>
  <span class="k">AVG</span><span class="p">(</span>
    <span class="k">CASE</span> <span class="k">WHEN</span> <span class="n">STRFTIME</span><span class="p">(</span><span class="s1">'%w'</span><span class="p">,</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span><span class="p">)</span> <span class="k">IN</span> <span class="p">(</span><span class="s1">'0'</span><span class="p">,</span> <span class="s1">'6'</span><span class="p">)</span> <span class="k">THEN</span> <span class="mi">1</span> <span class="k">ELSE</span> <span class="mi">0</span> <span class="k">END</span>
  <span class="p">)</span> <span class="k">AS</span> <span class="n">weekend_ratio</span><span class="p">,</span>
  <span class="k">AVG</span><span class="p">(</span>
    <span class="k">CASE</span> <span class="k">WHEN</span> <span class="n">STRFTIME</span><span class="p">(</span><span class="s1">'%w'</span><span class="p">,</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span><span class="p">)</span> <span class="k">IN</span> <span class="p">(</span><span class="s1">'0'</span><span class="p">,</span> <span class="s1">'6'</span><span class="p">)</span> <span class="k">THEN</span> <span class="mi">0</span> <span class="k">ELSE</span> <span class="mi">1</span> <span class="k">END</span>
  <span class="p">)</span> <span class="k">AS</span> <span class="n">weekday_ratio</span><span class="p">,</span>
  <span class="k">count</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">AS</span> <span class="n">total</span>
<span class="k">FROM</span> <span class="n">collisions</span> <span class="k">AS</span> <span class="k">c</span>
<span class="k">LEFT</span> <span class="k">JOIN</span> <span class="n">parties</span> <span class="k">AS</span> <span class="n">p</span>
  <span class="k">ON</span> <span class="k">c</span><span class="p">.</span><span class="n">case_id</span> <span class="o">=</span> <span class="n">p</span><span class="p">.</span><span class="n">case_id</span>
<span class="k">GROUP</span> <span class="k">BY</span> <span class="n">make</span>
<span class="k">HAVING</span> <span class="n">total</span> <span class="o">&gt;=</span> <span class="mi">10000</span>
<span class="k">ORDER</span> <span class="k">BY</span> <span class="n">weekday_ratio</span> <span class="k">DESC</span>
<span class="k">LIMIT</span> <span class="mi">1</span>
</code></pre></div>  </div>

  <p>Then I would copy and paste this but replace the <code class="language-plaintext highlighter-rouge">weekend_ratio</code> with a
<code class="language-plaintext highlighter-rouge">weekday_ratio</code>. It isn’t as “elegant” because we have to duplicate code, but
it is easy to write.</p>

  <p>Combining the queries is possible. To do so I first use a sub-query to do the
aggregation. A <code class="language-plaintext highlighter-rouge">WITH</code> clause keeps it tidy so we don’t have to copy/paste the
sub-query twice. I use <code class="language-plaintext highlighter-rouge">HAVING</code> to filter out makes with too few collisions;
it has to be <code class="language-plaintext highlighter-rouge">HAVING</code> and not <code class="language-plaintext highlighter-rouge">WHERE</code> because it filters <strong>after</strong> the
aggregation.</p>

  <p>I then construct two queries that read from the sub-query to select the
highest row for the weekend and weekdays. I <code class="language-plaintext highlighter-rouge">UNION</code> the two queries together
so we end up with a single table containing our results. The double select is
to allow the <code class="language-plaintext highlighter-rouge">ORDER BY</code> before the <code class="language-plaintext highlighter-rouge">UNION</code>.</p>

  <p>A note: for complicated queries like this one there are always many ways to do
it. I’d love to hear how you got it to work!</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">WITH</span> <span class="n">counter</span> <span class="k">AS</span> <span class="p">(</span>
  <span class="k">SELECT</span>
    <span class="n">p</span><span class="p">.</span><span class="n">vehicle_make</span> <span class="k">AS</span> <span class="n">make</span><span class="p">,</span>
    <span class="k">AVG</span><span class="p">(</span>
      <span class="k">CASE</span> <span class="k">WHEN</span> <span class="n">STRFTIME</span><span class="p">(</span><span class="s1">'%w'</span><span class="p">,</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span><span class="p">)</span> <span class="k">IN</span> <span class="p">(</span><span class="s1">'0'</span><span class="p">,</span> <span class="s1">'6'</span><span class="p">)</span> <span class="k">THEN</span> <span class="mi">1</span> <span class="k">ELSE</span> <span class="mi">0</span> <span class="k">END</span>
    <span class="p">)</span> <span class="k">AS</span> <span class="n">weekend_fraction</span><span class="p">,</span>
    <span class="k">AVG</span><span class="p">(</span>
      <span class="k">CASE</span> <span class="k">WHEN</span> <span class="n">STRFTIME</span><span class="p">(</span><span class="s1">'%w'</span><span class="p">,</span> <span class="k">c</span><span class="p">.</span><span class="n">collision_date</span><span class="p">)</span> <span class="k">IN</span> <span class="p">(</span><span class="s1">'0'</span><span class="p">,</span> <span class="s1">'6'</span><span class="p">)</span> <span class="k">THEN</span> <span class="mi">0</span> <span class="k">ELSE</span> <span class="mi">1</span> <span class="k">END</span>
    <span class="p">)</span> <span class="k">AS</span> <span class="n">weekday_fraction</span><span class="p">,</span>
    <span class="k">count</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">AS</span> <span class="n">total</span>
  <span class="k">FROM</span> <span class="n">collisions</span> <span class="k">AS</span> <span class="k">c</span>
  <span class="k">LEFT</span> <span class="k">JOIN</span> <span class="n">parties</span> <span class="k">AS</span> <span class="n">p</span>
    <span class="k">ON</span> <span class="k">c</span><span class="p">.</span><span class="n">case_id</span> <span class="o">=</span> <span class="n">p</span><span class="p">.</span><span class="n">case_id</span>
  <span class="k">GROUP</span> <span class="k">BY</span> <span class="n">make</span>
  <span class="k">HAVING</span> <span class="n">total</span> <span class="o">&gt;=</span> <span class="mi">10000</span>
<span class="p">)</span>

<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="p">(</span>
  <span class="k">SELECT</span>
    <span class="o">*</span>
  <span class="k">FROM</span> <span class="n">counter</span>
  <span class="k">ORDER</span> <span class="k">BY</span> <span class="n">weekend_fraction</span> <span class="k">DESC</span>
  <span class="k">LIMIT</span> <span class="mi">1</span>
<span class="p">)</span>

<span class="k">UNION</span>

<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="p">(</span>
  <span class="k">SELECT</span>
    <span class="o">*</span>
  <span class="k">FROM</span> <span class="n">counter</span>
  <span class="k">ORDER</span> <span class="k">BY</span> <span class="n">weekday_fraction</span> <span class="k">DESC</span>
  <span class="k">LIMIT</span> <span class="mi">1</span>
<span class="p">)</span>
</code></pre></div>  </div>

  <p>Which yields:</p>

  <table>
    <thead>
      <tr>
        <th style="text-align: left">make</th>
        <th style="text-align: right">weekend_fraction</th>
        <th style="text-align: right">weekday_fraction</th>
        <th style="text-align: right">total</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td style="text-align: left">HARLEY-DAVIDSON</td>
        <td style="text-align: right">0.385</td>
        <td style="text-align: right">0.614</td>
        <td style="text-align: right">49,602</td>
      </tr>
      <tr>
        <td style="text-align: left">PETERBILT</td>
        <td style="text-align: right">0.092</td>
        <td style="text-align: right">0.908</td>
        <td style="text-align: right">70,579</td>
      </tr>
    </tbody>
  </table>

  <p>These results makes sense, Peterbilt is a commercial truck manufacturer which
you expect to be driven for work. Harley-Davidson makes iconic motorcycles
that people ride for fun on the weekend with their friends.</p>

</details>

<h3 id="how-many-different-values-represent-toyota-in-the-parties-database-how-would-you-go-about-correcting-for-this">How many different values represent “Toyota” in the Parties database? How would you go about correcting for this?</h3>

<details>

  <summary>
    <p>Data is <strong><em>never</em></strong> as clean as you would hope,  and this applies even to the
<a href="/blog/switrs-sqlite-hosted-dataset/">curated SWITRS dataset</a>. How many different ways does
“Toyota” show up?</p>

    <p>What steps would you take to fix this problem?</p>

  </summary>

  <p>This is a case where there is no <em>right</em> answer. You can get a more and more
correct answer as you spend more time, but at some point you have to decide it
is good enough.</p>

  <p>The first step is to figure out what values might represent Toyota. I do that
with a few simple <code class="language-plaintext highlighter-rouge">LIKE</code> filters:</p>

  <div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span>
  <span class="n">vehicle_make</span><span class="p">,</span>
  <span class="k">COUNT</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span> <span class="k">AS</span> <span class="n">number_seen</span>
<span class="k">FROM</span> <span class="n">parties</span>
<span class="k">WHERE</span> <span class="k">LOWER</span><span class="p">(</span><span class="n">vehicle_make</span><span class="p">)</span> <span class="o">=</span> <span class="s1">'toyota'</span>
  <span class="k">OR</span> <span class="k">LOWER</span><span class="p">(</span><span class="n">vehicle_make</span><span class="p">)</span> <span class="k">LIKE</span> <span class="s1">'toy%'</span>
  <span class="k">OR</span> <span class="k">LOWER</span><span class="p">(</span><span class="n">vehicle_make</span><span class="p">)</span> <span class="k">LIKE</span> <span class="s1">'ty%'</span>
<span class="k">GROUP</span> <span class="k">BY</span> <span class="n">vehicle_make</span>
<span class="k">ORDER</span> <span class="k">BY</span> <span class="n">number_seen</span> <span class="k">DESC</span>
</code></pre></div>  </div>

  <p>Which gives us this table (truncated):</p>

  <div class="low-width-table" style="max-width: 20%">

    <table>
      <thead>
        <tr>
          <th style="text-align: left">vehicle_make</th>
          <th style="text-align: right">number_seen</th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td style="text-align: left">TOYOTA</td>
          <td style="text-align: right">2,374,621</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYO</td>
          <td style="text-align: right">166,209</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYT</td>
          <td style="text-align: right">146,746</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOT</td>
          <td style="text-align: right">2823</td>
        </tr>
        <tr>
          <td style="text-align: left">TOY</td>
          <td style="text-align: right">2262</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYTA</td>
          <td style="text-align: right">246</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOTA/</td>
          <td style="text-align: right">181</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYTO</td>
          <td style="text-align: right">84</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYTOA</td>
          <td style="text-align: right">71</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOYA</td>
          <td style="text-align: right">66</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYT.</td>
          <td style="text-align: right">65</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYA</td>
          <td style="text-align: right">51</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYTOTA</td>
          <td style="text-align: right">45</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOA</td>
          <td style="text-align: right">43</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYO /</td>
          <td style="text-align: right">39</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYT /</td>
          <td style="text-align: right">17</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYT/</td>
          <td style="text-align: right">14</td>
        </tr>
        <tr>
          <td style="text-align: left">TYMCO</td>
          <td style="text-align: right">13</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOTO</td>
          <td style="text-align: right">10</td>
        </tr>
        <tr>
          <td style="text-align: left">TOY0</td>
          <td style="text-align: right">10</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOYTA</td>
          <td style="text-align: right">7</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYTT</td>
          <td style="text-align: right">6</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOY</td>
          <td style="text-align: right">6</td>
        </tr>
        <tr>
          <td style="text-align: left">TOYOTS</td>
          <td style="text-align: right">5</td>
        </tr>
        <tr>
          <td style="text-align: left">TYOTA</td>
          <td style="text-align: right">4</td>
        </tr>
        <tr>
          <td style="text-align: left">…</td>
          <td style="text-align: right">…</td>
        </tr>
      </tbody>
    </table>

  </div>

  <p>Most of those look like they mean Toyota, although Tymco is a different
company that makes street sweepers.</p>

  <p>Here is how I would handle this issue: the top 5 make up the vast majority of
entries. I would fix those by hand and move on. More generally it seems that
makes are represented mostly by their name or a four-letter abbreviation. It
wouldn’t be too hard to detect and fix these for the most common makes.</p>

</details>

<h2 id="solutions">Solutions</h2>

<p>So that’s it! I hope it was useful and you learned something!</p>

<p>Here are my notebooks with the solutions:</p>

<ul>
  <li>
    <p>The <a href="/files/interview-prep/Interview%20Prep%20SQL%20Solutions.ipynb">SQL solution notebook</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/interview-prep/Interview%20Prep%20SQL%20Solutions.ipynb">Rendered on
Github</a>)</p>
  </li>
  <li>
    <p>The <a href="/files/interview-prep/Interview%20Prep%20Python%20Solutions.ipynb">Python/Pandas solution notebook</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/interview-prep/Interview%20Prep%20Python%20Solutions.ipynb">Rendered on
Github</a>)</p>
  </li>
</ul>

<p>A special thanks to <a href="https://github.com/quynhneo"><strong>Quynh M. Nguyen</strong></a> who came up with some simplifications for my queries!</p>

<p>Let me know if you find any more elegant solutions!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="interviewing" />
        
          <category term="interview-prep" />
        
      

      

      
      
        <summary type="html"><![CDATA[I often get asked how to practice data science interviews, so here is a practice dataset with a set of questions to answer. Good luck!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/interview-prep/food_conservation_workers_at_comstock_hall_cornell_1917.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/interview-prep/food_conservation_workers_at_comstock_hall_cornell_1917.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Introducing the SWITRS SQLite Hosted Dataset</title>
      <link href="https://alexgude.com/blog/switrs-sqlite-hosted-dataset/" rel="alternate" type="text/html" title="Introducing the SWITRS SQLite Hosted Dataset" />
      <published>2020-11-24T00:00:00-08:00</published>
      <updated>2020-11-24T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_sqlite_hosted_dataset</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-sqlite-hosted-dataset/"><![CDATA[<p>The State of California maintains a database of traffic collisions called the
<a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">Statewide Integrated Traffic Records System (SWITRS)</a>. I have made
extensive use of the data, including:</p>

<ul>
  <li>
    <p>Finding out when <a href="/blog/switrs-crashes-by-date/">automobiles</a>, <a href="/blog/switrs-motorcycle-crashes-by-date/">motorcycles</a>, and
<a href="/blog/switrs-bicycle-crashes-by-date/">bicycles</a> crash.</p>
  </li>
  <li>
    <p>Quantifying the dangers of <a href="/blog/switrs-daylight-saving-time-accidents/">daylight saving time</a> and the <a href="/blog/switrs-daylight-saving-time-end-accidents/">end of
daylight saving time</a>.</p>
  </li>
</ul>

<p>I even maintain a <a href="/blog/switrs-to-sqlite/">helpful script</a> to convert the messy CSV files
California will let you download into a clean <a href="https://en.wikipedia.org/wiki/SQLite">SQLite database</a>.</p>

<p>But requesting the data is painful; you have to set up an account with the
state, submit your request via a rough web-form, and wait for them to compile
it. Worse, in the last few years California has <em>limited the data to 2008 and
later!</em></p>

<p>Luckily I have saved all the data I’ve requested which goes back to January
1st, 2001. So I resolved to make the data easily available to everyone.</p>

<h2 id="the-switrs-hosted-dataset">The SWITRS Hosted Dataset</h2>

<p>I have combined all of my data requests into one SQLite database. You no
longer have to worry about requesting the data or using my script to clean it
up since I have done all that work for you.</p>

<p>You can <strong>download the database</strong> from either <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs"><strong>Kaggle</strong> (requires
account)</a> or <a href="https://zenodo.org/record/4284843"><strong>Zenodo</strong> (no account required!)</a> and get
right to work! I have even included a <a href="https://www.kaggle.com/alexgude/starter-california-traffic-collisions-from-switrs">demo notebook on Kaggle</a> so
you can jump right in!</p>

<p>The dataset also has its own <a href="https://en.wikipedia.org/wiki/Digital_object_identifier">DOI</a>: <a href="https://www.doi.org/10.34740/kaggle/dsv/1671261">10.34740/kaggle/dsv/1671261</a></p>

<p>Read on for an example of how to use the dataset and an explanation of how I
created it.</p>

<h3 id="data-merging">Data Merging</h3>

<p>I have saved four copies of the data, requested in 2016, 2017, 2018, and 2020.
The first three copies have data from 2001 until their request date, while the
2020 dataset only covers 2008–2020 due to the new limit California
instituted. To create the hosted dataset I had to merge these four datasets.
There were two main challenges:</p>

<ol>
  <li>
    <p>Each dataset contains three tables: collision records, party records, and
victim records; but <em>only</em> the collision records table contains a
<a href="https://en.wikipedia.org/wiki/Primary_key"><strong>primary key</strong></a>. That key is the <code class="language-plaintext highlighter-rouge">case_id</code>.</p>
  </li>
  <li>
    <p>The records are occasionally updated after the fact, but again only the
collision records table has a column (<code class="language-plaintext highlighter-rouge">process_date</code>) indicating when the
record was last modified.</p>
  </li>
</ol>

<p>I made the following assumptions when merging the datasets:</p>

<ul>
  <li>
    <p>The collision records table from the more recent dataset was correct when
there was a conflict.</p>
  </li>
  <li>
    <p>The party records and victim records corresponding to that collision record
were also the most correct.</p>
  </li>
</ul>

<p>These assumptions allowed me to write out the following join logic to create
the hosted set. First I selected <code class="language-plaintext highlighter-rouge">case_id</code> from each copy of the data,
preferring the newer ones:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Select all from 2020</span>
<span class="k">CREATE</span> <span class="k">TABLE</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">case_ids</span> <span class="k">AS</span>
<span class="k">SELECT</span> <span class="n">case_id</span><span class="p">,</span> <span class="s1">'2020'</span> <span class="k">AS</span> <span class="n">db_year</span>
<span class="k">FROM</span> <span class="n">db20</span><span class="p">.</span><span class="n">collision</span><span class="p">;</span>

<span class="c1">-- Now add the rows that don't match from earlier databases, in</span>
<span class="c1">-- reverse chronological order so that the newer rows are not</span>
<span class="c1">-- overwritten.</span>
<span class="k">INSERT</span> <span class="k">INTO</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">case_ids</span>
<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="p">(</span>
    <span class="k">SELECT</span> <span class="n">older</span><span class="p">.</span><span class="n">case_id</span><span class="p">,</span> <span class="s1">'2018'</span>
    <span class="k">FROM</span> <span class="n">db18</span><span class="p">.</span><span class="n">collision</span> <span class="k">AS</span> <span class="n">older</span>
    <span class="k">LEFT</span> <span class="k">JOIN</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">case_ids</span> <span class="k">AS</span> <span class="n">prime</span>
    <span class="k">ON</span> <span class="n">prime</span><span class="p">.</span><span class="n">case_id</span> <span class="o">=</span> <span class="n">older</span><span class="p">.</span><span class="n">case_id</span>
    <span class="k">WHERE</span> <span class="n">prime</span><span class="p">.</span><span class="n">case_id</span> <span class="k">IS</span> <span class="k">NULL</span>
<span class="p">);</span>

<span class="c1">-- and the same for 2017 and 2016</span>
</code></pre></div></div>

<p>Then I selected the rows from the collision records, part records, and victim
records that matched for each year:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">CREATE</span> <span class="k">TABLE</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">collision</span> <span class="k">AS</span>
<span class="k">SELECT</span> <span class="o">*</span>
<span class="k">FROM</span> <span class="n">db20</span><span class="p">.</span><span class="n">collision</span><span class="p">;</span>

<span class="k">INSERT</span> <span class="k">INTO</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">collision</span>
<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="p">(</span>
    <span class="k">SELECT</span> <span class="n">col</span><span class="p">.</span><span class="o">*</span>
    <span class="k">FROM</span> <span class="n">db18</span><span class="p">.</span><span class="n">collision</span> <span class="k">AS</span> <span class="n">col</span>
    <span class="k">INNER</span> <span class="k">JOIN</span> <span class="n">outputdb</span><span class="p">.</span><span class="n">case_ids</span> <span class="k">AS</span> <span class="n">ids</span>
    <span class="k">ON</span> <span class="n">ids</span><span class="p">.</span><span class="n">case_id</span> <span class="o">=</span> <span class="n">col</span><span class="p">.</span><span class="n">case_id</span>
    <span class="k">WHERE</span> <span class="n">ids</span><span class="p">.</span><span class="n">db_year</span> <span class="o">=</span> <span class="s1">'2018'</span>
<span class="p">);</span>

<span class="c1">-- and similarly for 2017 and 2016, and</span>
<span class="c1">-- for party records and victim records</span>
</code></pre></div></div>

<p>The <a href="https://github.com/agude/SWITRS-to-SQLite/blob/master/scripts/combine_databases.sql">script to do this is here</a>.</p>

<h3 id="using-the-dataset">Using the dataset</h3>

<p>Using the hosted dataset, it is simple to reproduce the work I did when I
announced the data converter script: <a href="/blog/switrs-to-sqlite/#crash-mapping-example">plotting the location of all crashes in
California</a>.</p>

<p>Just download the data, unzip it, and run the <a href="/files/switrs-dataset/Mapping%20California%20Crashes%202001%20to%202020.ipynb">notebook</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-dataset/Mapping%20California%20Crashes%202001%20to%202020.ipynb">rendered
on Github</a>). This will produce the following plot:</p>

<p><a href="/files/switrs-dataset/2001-2020_california_traffic_collisions_map.png"><img src="/files/switrs-dataset/2001-2020_california_traffic_collisions_map.png" alt="A map of the location of all the crashes in the state of California from
2001 to 2020" /></a></p>

<p>I hope this <a href="https://www.kaggle.com/alexgude/california-traffic-collision-data-from-switrs">hosted dataset</a> makes working with the data fast and
easy. If you make something, I’d love to see it! Send it to me on BlueSky:
<a href="https://bsky.app/profile/alexgude.com">@alexgude.com</a></p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[California traffic collision data has been hard to get, that's why I am now curating and hosting it! Come take a look!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-dataset/tram_auto_crash_in_1957_frederiksplein_amsterdam.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-dataset/tram_auto_crash_in_1957_frederiksplein_amsterdam.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Plotting the 2020 Tour de France</title>
      <link href="https://alexgude.com/blog/2020-tour-de-france-plot/" rel="alternate" type="text/html" title="Plotting the 2020 Tour de France" />
      <published>2020-10-16T00:00:00-07:00</published>
      <updated>2020-10-16T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/2020_tour_de_france_plot</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/2020-tour-de-france-plot/"><![CDATA[<p>The <a href="https://en.wikipedia.org/wiki/2020_Tour_de_France">Tour de France</a> was postponed by <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">the pandemic</a> this year,
but finally kicked off in late August. Although there were worries that the
race would have to be stopped in the middle, it made it all the way to the
final sprint on the Champs-Élysées in Paris. In this post, just like <a href="/blog/2019-tour-de-france-plot/">last
year’s</a>, I will use plots to explore how the Tour unfolded.</p>

<p>The code that generated the plots can be found <a href="/files/tour-de-france//Tour%20de%20France%202020%20Plot.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/tour-de-france//Tour%20de%20France%202020%20Plot.ipynb">rendered on
Github</a>). The data <a href="/files/tour-de-france//2020-tdf-dataframe.json">is here</a>.</p>

<h2 id="the-race-for-yellow">The Race for Yellow</h2>

<p>The most prestigious award at the Tour is the <a href="https://en.wikipedia.org/wiki/General_classification_in_the_Tour_de_France">yellow jersey</a>, which
is awarded to the rider with the lowest combined time across all 21 stages of
the race. <a href="https://en.wikipedia.org/wiki/Egan_Bernal">Egan Bernal</a> was the favorite going into this year as he
had won last year’s race. His team, <a href="https://en.wikipedia.org/wiki/Ineos_Grenadiers">Ineos</a>, had also decided to
dedicate all of their resources to him and left former winner and previous
co-leads <a href="https://en.wikipedia.org/wiki/Chris_Froome">Chris Froome</a> and <a href="https://en.wikipedia.org/wiki/Geraint_Thomas">Geraint Thomas</a> off the roster.</p>

<p><a href="https://en.wikipedia.org/wiki/Primo%C5%BE_Rogli%C4%8D">Primož Roglič</a> was another favorite. He had won last year’s <a href="https://en.wikipedia.org/wiki/2019_Vuelta_a_Espa%C3%B1a">Vuelta a
España</a>, taken 4th in a previous Tour, and his team,
<a href="https://en.wikipedia.org/wiki/Team_Jumbo%E2%80%93Visma">Jumbo–Visma</a>, included a star-studded support roster.</p>

<p>After three weeks of racing, here is how the top five riders fared through the
stages:</p>

<p><a href="/files/tour-de-france//2020_tour_de_france_top_5.svg"><img src="/files/tour-de-france//2020_tour_de_france_top_5.svg" alt="A line plot showing how far behind the leader each top-finishing rider was
after each stage of the 2020 Tour de France." /></a></p>

<p>Stage 7 stands out in this plot. Although large time gaps normally occur
during mountain finishes, this completely flat stage shook up the race for
yellow. Instead of a steep climb, strong winds split the <a href="https://en.wikipedia.org/wiki/Peloton">peloton</a> in
two and several top riders were stuck in the chasing group where they lost
1′21″.</p>

<p>After leading for most of the race, Roglič lost nearly a minute to <a href="https://en.wikipedia.org/wiki/Tadej_Poga%C4%8Dar">Tadej
Pogačar</a>, a young Slovenian riding his first Tour ever, on the
penultimate stage. Roglič had defended the jersey since stage 9, possibly as
part of a strategy to take the jersey early in case the race had to be
canceled midway through.</p>

<p>But the long defense left Roglič vulnerable. In a ride that caused 17-time
Tour rider <a href="https://en.wikipedia.org/wiki/George_Hincapie">George Hincapie</a> to declare it <a href="https://web.archive.org/web/20230315051620/https://wedu.team/themove/2020-tour-de-france-stage-20">“the greatest Tour I
have ever seen”</a>, Pogačar stormed back on <a href="https://en.wikipedia.org/wiki/La_Planche_des_Belles_Filles">La Planche des Belles
Filles</a>, taking first on the stage by 1′21″. Every other top rider
lost time on the stage as well, even <a href="https://en.wikipedia.org/wiki/Richie_Porte">Richie Porte</a> who came in third
for the stage and knocked <a href="https://en.wikipedia.org/wiki/Miguel_%C3%81ngel_L%C3%B3pez_(cyclist)">Miguel Ángel López</a> out of the top three
overall.</p>

<h3 id="disappointing-results">Disappointing Results</h3>

<p>Many riders set their sights on the yellow jersey but ultimately fall short.
Crashes, illness, and simply not being in form can drag down even top riders.
Here are the riders who went for the glory but could not keep it up for the
full three weeks:</p>

<p><a href="/files/tour-de-france//2020_tour_de_france_underperforming.svg"><img src="/files/tour-de-france//2020_tour_de_france_underperforming.svg" alt="A line plot showing how some of the under-performing riders fell in the
rankings." /></a></p>

<p>Notice that the y-axis now extends to over two hours behind the leader, not
the minutes behind in the first chart.</p>

<p>It might be hard to call a top 6 finish a disappointment, but for Lopez is
was. He was on the podium in 3rd place when he started stage 20, but he lost
over 6 minutes in a disastrous time trial.</p>

<p><a href="https://en.wikipedia.org/wiki/Guillaume_Martin">Guillaume Martin</a> finished 11th, his highest ever place, but he had
been in the top three for much of the early race, keeping up with favorites
Bernal and Roglič. He lost time during stage 13 after holding strong during
the first real test in the Pyrenees.</p>

<p>Both Bernal—last year’s winner—and <a href="https://en.wikipedia.org/wiki/Nairo_Quintana">Nario Quintana</a>—two time
runner up to Chris Froome—defended well in the early mountains but lost time
in the high <a href="https://en.wikipedia.org/wiki/Massif_Central">Massif Central</a>. They were suffering from injuries incurred
during crashes earlier in the race. In a controversial move, Bernal withdrew
from the race after he lost time.<sup style="anchor-name:--fnref-sportsmanship" id="fnref:sportsmanship"><a href="#fn:sportsmanship" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Quintana fought on and
finished in Paris, but lost lots of time in the Alps.</p>

<p><a href="https://en.wikipedia.org/wiki/Thibaut_Pinot">Thibaut Pinot</a> crashed on stage 1. <a href="https://en.wikipedia.org/wiki/Emanuel_Buchmann">Emanuel Buchmann</a>
crashed in a previous race and his ability to start was in question. Both lost
time in the first mountains and never recovered, but nevertheless stayed in
the race through the end.</p>

<h2 id="the-rest-of-the-race">The Rest of the Race</h2>

<p>More than 100 riders finished the Tour, but most of them were not competing
for the yellow jersey. Here are the paths taken by all 146 riders who finished
in Paris:</p>

<p><a href="/files/tour-de-france//2020_tour_de_france.svg"><img src="/files/tour-de-france//2020_tour_de_france.svg" alt="A line plot showing how far behind the leader every rider was for each
stage." /></a></p>

<p>Buchmann, the lowest placed rider in our previous plot, is actually ahead of
most of the riders! We can also see the race started out tough, probably due
to the chance that it might be canceled after the first rest day, with large
time gaps opening up even before the first mountains.</p>

<p>The latter half of the second week was also tough with the hilly Massif
Central and mountainous Alps. By the last few stages, the time gaps were
pretty much set and most riders maintained their relative positions.</p>

<p>Two riders of note are <a href="https://en.wikipedia.org/wiki/Peter_Sagan">Peter Sagan</a> and <a href="https://en.wikipedia.org/wiki/Sam_Bennett_(cyclist)">Sam Bennett</a>, who
were competing for the <a href="https://en.wikipedia.org/wiki/Points_classification_in_the_Tour_de_France">green jersey</a>. Both of them saved their energy,
and hence lost time, on mountainous stages so they could give it their all in
the sprints. Even though Sagan finished about an hour ahead of Bennett, he
lost the Jersey. Bennett had done a better job of managing his energy and
using it where it counted.</p>

<p>Finally, <a href="https://en.wikipedia.org/wiki/Roger_Kluge">Roger Kluge</a> won the <a href="https://en.wikipedia.org/wiki/Lanterne_rouge">lanterne rouge</a>, finishing
six hours behind Pogačar. His job in the race had been to escort his team’s
sprinter, <a href="https://en.wikipedia.org/wiki/Caleb_Ewan">Caleb Ewan</a>, through the race. This often meant falling back
on climbs and waiting for Ewan so they could tackle the mountains together and
avoid being cut for being too slow.</p>

<p>This year’s Tour was unique due to needing to adjust to the COVID pandemic,
but it turned out to be one of the most exciting races in the history of the
sport! And what’s more, it gave us some much-needed entertainment during these
dark times.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:sportsmanship">

      <p>Ineos said Bernal dropped out to “focus on recovery”, but many fans felt
that Bernal—who had won last year, placed as high as second this year,
and worn the <a href="https://en.wikipedia.org/wiki/White_jersey">white jersey</a> for the best young rider—was
abandoning the most prestigious race of the season to avoid embarrassment
at the hands of his opponents. <a href="#fnref:sportsmanship" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[The Tour de France is a race decided by mere minutes; to see exactly how those minutes were earned, read on for my plots!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_second.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932_second.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Data Science Interviews During the 2020 Pandemic</title>
      <link href="https://alexgude.com/blog/interviewing-for-data-science-positions-in-2020/" rel="alternate" type="text/html" title="Data Science Interviews During the 2020 Pandemic" />
      <published>2020-09-21T00:00:00-07:00</published>
      <updated>2020-09-21T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/interviewing_for_data_science_positions_in_2020</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/interviewing-for-data-science-positions-in-2020/"><![CDATA[<p>I started looking for a job this year because Intuit cut my old position as
part of <a href="https://en.wikipedia.org/wiki/COVID-19_pandemic">COVID-related</a> layoffs. Consequently, I spent the last three
months preparing for and participating in job interviews.</p>

<p>The interview process was different from when I first interviewed as a young
data scientist right out of <a href="/blog/should-i-go-to-insight/">Insight</a>, and even different from my
more recent interview experiences in 2017.</p>

<h2 id="observations">Observations</h2>

<p>All of the interviews this year followed <a href="/blog/interviews-respect-time/">the structure I outlined in my post
on not wasting a candidate’s time</a>, which I appreciated. To summarize,
that structure is:</p>

<ul>
  <li>
    <p><strong>Prescreen</strong>: A resume or recruiter screen of the candidate, often offline.</p>
  </li>
  <li>
    <p><strong>Technical Screen</strong>: Either a call with a team member or a take-home
assignment to assess the candidate’s technical skills.</p>
  </li>
  <li>
    <p><strong>On-site Interview Loop</strong>: An all-day interview with multiple team members
and the hiring manager.</p>
  </li>
</ul>

<p>Although all of the companies structured the interviews well, several failed
at the second area I emphasize in my post: <strong>Communication</strong>. One company took
an entire month to convey feedback from an interview step while another did
not get back to me about my on-site until I had emailed them several times.</p>

<h3 id="salary-negotiation">Salary Negotiation</h3>

<p>The first conversation with every recruiter in 2017 involved talking about
present salary and future expectations. In 2020, not a single recruiter asked
about compensation until after they had made a verbal offer. This surprised me
because waiting to open negotiations until that late in the process offers the
candidate an advantage as they now know that the company has strong interest
in them. Previously I have tried to delay this conversation as long as
possible as part of <a href="/blog/data-science-asking-for-more-money/">my negotiation strategy</a>, but this time
the recruiters did it for me!</p>

<p>This behavior might be explained by <a href="https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?sectionNum=432.3&amp;lawCode=LAB">California’s new law which bans using
salary history to determine an offer</a>, but it specifically <em>does
not</em> ban asking about salary expectations. It might be that the more senior
roles I’m interviewing for now are hard enough to find candidates for that the
companies don’t want to reject anyone without a chance to get the candidate
<a href="https://en.wikipedia.org/wiki/Loss_aversion">committed to the role</a>.</p>

<h3 id="technical-screens">Technical Screens</h3>

<p>In 2015 and 2017 I had real trouble with the technical screens. Often they
were stereotypical <em>engineering</em> interviews where I was asked to do something
difficult but irrelevant, like <a href="https://twitter.com/mxcl/status/608682016205344768">invert a binary search tree</a>. I was
able to solve these problems when I had seen them before in my studies or
could come up with the <em>“trick”</em> on the fly. I passed about half of my
technical screens.</p>

<p>In 2020, my experience was vastly different. Only one screen was even close to
“invert this BST”, and that was a mistake where they later admitted that I was
given the <em>software engineering</em> technical screen by mistake.</p>

<p>All of the other screens involved reasonable questions that would come up in a
data scientist’s daily work, like <a href="/blog/data-science-interview-prep-data-manipulation/">manipulating a
dataset</a>, calculating some features, or implementing
really simple algorithms or metrics. With these more applied questions, I
passed all of my technical screens!</p>

<h3 id="virtual-on-sites">Virtual On-Sites</h3>

<p>In 2017 on-sites were on site! In 2020 they are done via video conferencing. I
thought virtual on-sites would be less draining, but I actually felt even more
exhausted after them. However, I still felt I was able to connect on a
personal level with the interviewers, despite not being in the same place.</p>

<h4 id="whiteboard-coding">Whiteboard Coding</h4>

<p>There were fewer coding problems during the on-sites than previously and all
of them were done in an online editor instead of on a whiteboard. This worked
great!</p>

<p>I found myself looking forward to the coding challenges because, with the
improvement of coding on <strong>an actual computer</strong>, they were a nice break from
the other interviews. Just like the technical screens, these questions were
all directly applicable to the work I would be doing.</p>

<h4 id="open-ended-problems-and-behavioral-questions">Open-ended Problems and Behavioral Questions</h4>

<p>This time there were more open-ended interviews that dug into some problem the
business had (for example “How would you help us filter spam?”—similar to
the <a href="/blog/data-science-interview-prep-case-study/">machine learning case studies I’ve written about for interview
prep</a>) or went really deep exploring a project I had worked on
previously. Although I was asked some of these questions during my previous
years of interviewing they felt much more effective in the virtual format,
perhaps because the lack of a whiteboard made it so the interviewer and I had
to have a conversation instead of me giving a lecture.</p>

<p>This round of interviews was the first time I was asked <a href="https://en.wikipedia.org/wiki/Job_interview#Behavioral_interview_questions">behavioral
questions</a>. They were present in three of the five on-sites, including
one company that had 90 minutes of them!</p>

<h2 id="results">Results</h2>

<p>I applied to seven companies using internal referrals from my network. Here is
how I did during each round:</p>

<table>
  <thead>
    <tr>
      <th><strong>Company</strong></th>
      <th style="text-align: right">Prescreen</th>
      <th style="text-align: right">Technical Screen</th>
      <th style="text-align: right">On-Site</th>
      <th style="text-align: right">Offer</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>DocuSign</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:DarkBlue">Declined</span><sup style="anchor-name:--fnref-docusign" id="fnref:docusign"><a href="#fn:docusign" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td><strong>Grand Rounds</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:Red">Reject</span></td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td><strong>Salesforce</strong></td>
      <td style="text-align: right"><span style="color:Red">Reject</span></td>
      <td style="text-align: right">—</td>
      <td style="text-align: right">—</td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td><strong>Square</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Accepted</span></td>
    </tr>
    <tr>
      <td><strong>Stripe</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:Red">Reject</span></td>
      <td style="text-align: right">—</td>
    </tr>
    <tr>
      <td><strong>Twitch</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:DarkBlue">Declined</span></td>
    </tr>
    <tr>
      <td><strong>Twitter</strong></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:ForestGreen">Pass</span></td>
      <td style="text-align: right"><span style="color:Red">Reject</span></td>
      <td style="text-align: right">—</td>
    </tr>
  </tbody>
</table>

<p>I am very happy with how well the technical screens went this time around, as
mentioned above. I also feel good about the on-site to offer rate, although
more offers would always be better of course.</p>

<p>I felt three interviews went really well: Square, Twitch, and Twitter. I
thought I connected well with the teams and demonstrated that I had the skills
they were looking for. So I was disappointed to not get an offer from Twitter,
but excited for the offers from Square and Twitch.</p>

<p>The Stripe interview revealed the position to be less well aligned with my
skill set<sup style="anchor-name:--fnref-ab" id="fnref:ab"><a href="#fn:ab" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> and career aspirations than I’d hoped, so I think their
decision to not extend an offer was fair.</p>

<p>I generally do not apply to work for startups, for a variety of reasons both
personal and <a href="https://every.to/napkin-math/you-probably-shouldn-t-work-at-a-startup-9387b632-345c-4a22-bac0-3cb92f0eecf1">financial</a>.<sup style="anchor-name:--fnref-sense" id="fnref:sense"><a href="#fn:sense" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> However, I agreed to interview at
Grand Rounds after a friend reached out because I did not want to turn down
any opportunities. In the end we all knew it was a bad fit.</p>

<p>So, after all that, I’m excited to get to work at Square and even more excited
to be done interviewing during a pandemic!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:docusign">
      <p>DocuSign and I were unable to schedule an on-site before my other offers would have expired. <a href="#fnref:docusign" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ab">
      <p>The role was a little heavier on <a href="https://www.projectpro.io/article/type-a-data-scientist-vs-type-b-data-scientist/194">the “analytics” side of data science than the “building” side</a> which I prefer. <a href="#fnref:ab" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:sense">

      <p>The previous article I linked here has been purged from the
internet, so I found a replacement by Evan Armstrong. I’ve copied the
relevant part below:</p>

      <blockquote>
        <p>Say you have a 2% chance of picking a unicorn and being a member of
the founding team. That 2% is honestly way too high for most people,
and perhaps a bit low for others, but a good median to anchor on. $10M * 2% = $200K. And realistically you’re only going to get that this
tranche of equity every 3–4 years at most, so that’s a risk-adjusted
value of $50–$66k.</p>

        <p>In contrast, if you were to take a job at a tier 1 tech firm, such as
Facebook or Google, as a high-quality engineer your salary can be
$200–400k with an additional $100–250K in equity. You will receive the
best benefits package known to mankind, with massages, free food, the
finest gear money can buy, 401Ks, bonus programs, copious amounts of
vacation, and the list goes on. Your total compensation package of
salary + equity + benefits will be far higher versus a startup.</p>
      </blockquote>
      <p><a href="#fnref:sense" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="interviewing" />
        
      

      

      
      
        <summary type="html"><![CDATA[In the middle of the COVID-19 pandemic, I found myself looking for a data science job for the third time in my life. This post covers what I learned.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/2020-interviewing/nighhawks_crop.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/2020-interviewing/nighhawks_crop.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Python Patterns: Map and Filter</title>
      <link href="https://alexgude.com/blog/python-patterns-map-filter/" rel="alternate" type="text/html" title="Python Patterns: Map and Filter" />
      <published>2020-08-31T00:00:00-07:00</published>
      <updated>2020-08-31T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/python_patterns_map_filter</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-patterns-map-filter/"><![CDATA[<p>Computers are great at performing a simple action over and over again. A common
way to make them do such a task is to store data in a list and iterate over it
with a for loop, calling a function for each item.</p>

<p>But Python has some great functions to replace for loops, which I will cover
below after a quick example.</p>

<h2 id="playing-cards">Playing Cards</h2>

<p>Given a list of playing cards as tuples, like so:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">cards</span> <span class="o">=</span> <span class="p">[</span>
  <span class="p">(</span><span class="sh">"</span><span class="s">Spades</span><span class="sh">"</span><span class="p">,</span> <span class="mi">14</span><span class="p">),</span>
  <span class="p">(</span><span class="sh">"</span><span class="s">Diamonds</span><span class="sh">"</span><span class="p">,</span> <span class="mi">13</span><span class="p">),</span>
  <span class="p">(</span><span class="sh">"</span><span class="s">Hearts</span><span class="sh">"</span><span class="p">,</span> <span class="mi">2</span><span class="p">),</span>
  <span class="p">(</span><span class="sh">"</span><span class="s">Spades</span><span class="sh">"</span><span class="p">,</span> <span class="mi">8</span><span class="p">),</span>
  <span class="p">(</span><span class="sh">"</span><span class="s">Clubs</span><span class="sh">"</span><span class="p">,</span> <span class="mi">11</span><span class="p">),</span>
  <span class="p">...</span>  <span class="c1"># etc.
</span><span class="p">]</span>
</code></pre></div></div>

<p>We want to convert them to <code class="language-plaintext highlighter-rouge">PlayingCard</code> objects as <a href="/blog/python-patterns-enum/#playing-cards-with-enums">defined in my previous
post on <code class="language-plaintext highlighter-rouge">enums</code></a>. To do this, we need a function to convert a tuple into the
class:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">):</span>
  <span class="n">suit</span><span class="p">,</span> <span class="n">rank</span> <span class="o">=</span> <span class="n">card_tuple</span>

  <span class="n">card</span> <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span>
    <span class="nc">CardSuit</span><span class="p">(</span><span class="n">suit</span><span class="p">),</span>
    <span class="nc">CardRank</span><span class="p">(</span><span class="n">rank</span><span class="p">),</span>
  <span class="p">)</span>

  <span class="k">return</span> <span class="n">card</span>
</code></pre></div></div>

<p>This makes it easy to parse the list with a quick loop:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">new_cards</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">card_tuple</span> <span class="ow">in</span> <span class="n">cards</span><span class="p">:</span>
  <span class="n">new_card</span> <span class="o">=</span> <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">)</span>
  <span class="n">new_cards</span><span class="p">.</span><span class="nf">append</span><span class="p">(</span><span class="n">new_card</span><span class="p">)</span>
</code></pre></div></div>

<p>And we can even filter the cards so that we only keep hearts:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">just_hearts</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">card_tuple</span> <span class="ow">in</span> <span class="n">cards</span><span class="p">:</span>
  <span class="n">new_card</span> <span class="o">=</span> <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">)</span>
  <span class="k">if</span> <span class="n">new_card</span><span class="p">.</span><span class="n">suit</span> <span class="ow">is</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span><span class="p">:</span>
    <span class="n">just_hearts</span><span class="p">.</span><span class="nf">append</span><span class="p">(</span><span class="n">new_card</span><span class="p">)</span>
</code></pre></div></div>

<p>These code snippets are fine: short and clean, with not a lot that can go wrong. But I
love to replace custom code with Python built-ins whenever possible, because
they are fast, well tested, and concise. Python provides two functions that
can simplify even these already simple code fragments: <code class="language-plaintext highlighter-rouge">map()</code> and <code class="language-plaintext highlighter-rouge">filter()</code>.</p>

<h2 id="map">Map</h2>

<p>The <a href="https://docs.python.org/3.7/library/functions.html#map"><code class="language-plaintext highlighter-rouge">map()</code> function</a> replaces a for loop that calls a function on each
item of a list, just as we did in the above when making <code class="language-plaintext highlighter-rouge">PlayingCard</code> objects.
Here is how we could rewrite the above code using map:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">new_cards</span> <span class="o">=</span> <span class="nf">map</span><span class="p">(</span><span class="n">tuple_to_card</span><span class="p">,</span> <span class="n">cards</span><span class="p">)</span>
</code></pre></div></div>

<p>Simple!</p>

<p>If the function is not too complicated, it can be useful to define it inline
with a <a href="https://docs.python.org/3/reference/expressions.html#lambda"><code class="language-plaintext highlighter-rouge">lambda</code> function</a>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">get_card_rank</span> <span class="o">=</span> <span class="k">lambda</span> <span class="n">card_tuple</span><span class="p">:</span> <span class="n">card_tuple</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
<span class="n">ranks</span> <span class="o">=</span> <span class="nf">map</span><span class="p">(</span><span class="n">get_card_rank</span><span class="p">,</span> <span class="n">cards</span><span class="p">)</span>
</code></pre></div></div>

<p>But what if instead of tuples we had two lists: one of suits and one of ranks?
We could use the <a href="https://docs.python.org/3.7/library/functions.html#zip"><code class="language-plaintext highlighter-rouge">zip</code> function</a> to combine the two lists like a zipper:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">new_cards</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">card_tuple</span> <span class="ow">in</span> <span class="nf">zip</span><span class="p">(</span><span class="n">card_suits</span><span class="p">,</span> <span class="n">card_ranks</span><span class="p">):</span>
  <span class="n">new_card</span> <span class="o">=</span> <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">)</span>
  <span class="n">new_cards</span><span class="p">.</span><span class="nf">append</span><span class="p">(</span><span class="n">new_card</span><span class="p">)</span>
</code></pre></div></div>

<p>But map already allows us to do pairwise operations:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">new_cards</span> <span class="o">=</span> <span class="nf">map</span><span class="p">(</span><span class="n">tuple_to_card</span><span class="p">,</span> <span class="n">card_suits</span><span class="p">,</span> <span class="n">card_ranks</span><span class="p">)</span>
</code></pre></div></div>

<p>Of course, we could write our own map using <a href="https://docs.python.org/3/tutorial/datastructures.html#list-comprehensions">list comprehension</a>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">new_cards</span> <span class="o">=</span> <span class="p">[</span><span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">)</span> <span class="k">for</span> <span class="n">card_tuple</span> <span class="ow">in</span> <span class="n">cards</span><span class="p">]</span>
</code></pre></div></div>

<p>Which, perhaps, is a little more Pythonic.</p>

<h2 id="filter">Filter</h2>

<p>But how would we filter the list so that we only keep hearts, as in our second
example? We could wrap the <code class="language-plaintext highlighter-rouge">map</code> call in a for loop:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">just_hearts</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">card</span> <span class="ow">in</span> <span class="nf">map</span><span class="p">(</span><span class="n">tuple_to_card</span><span class="p">,</span> <span class="n">cards</span><span class="p">):</span>
  <span class="k">if</span> <span class="n">card</span><span class="p">.</span><span class="n">suit</span> <span class="ow">is</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span><span class="p">:</span>
    <span class="n">just_hearts</span><span class="p">.</span><span class="nf">append</span><span class="p">(</span><span class="n">card</span><span class="p">)</span>
</code></pre></div></div>

<p>But the <a href="https://docs.python.org/3.7/library/functions.html#filter"><code class="language-plaintext highlighter-rouge">filter()</code> function</a> does that for us! It takes a function
and an iterable and returns only the elements of the iterable that evaluate to
<code class="language-plaintext highlighter-rouge">True</code> when the function is called on them. This allows us to rewrite the
above as:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">is_a_heart</span> <span class="o">=</span> <span class="k">lambda</span> <span class="n">card</span><span class="p">:</span> <span class="n">card</span><span class="p">.</span><span class="n">suit</span> <span class="ow">is</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span>

<span class="n">just_hearts</span> <span class="o">=</span> <span class="nf">filter</span><span class="p">(</span>
  <span class="n">is_a_heart</span><span class="p">,</span>
  <span class="nf">map</span><span class="p">(</span><span class="n">tuple_to_card</span><span class="p">,</span> <span class="n">cards</span><span class="p">)</span>
<span class="p">)</span>
</code></pre></div></div>

<p>Of course, again, we could write this as a comprehension:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">just_hearts</span> <span class="o">=</span> <span class="p">[</span>
  <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">)</span> <span class="k">for</span> <span class="n">card_tuple</span> <span class="ow">in</span> <span class="n">cards</span>
  <span class="k">if</span> <span class="nf">tuple_to_card</span><span class="p">(</span><span class="n">card_tuple</span><span class="p">).</span><span class="n">suit</span> <span class="ow">is</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span>
<span class="p">]</span>
</code></pre></div></div>

<p>However, this is not as readable as the <code class="language-plaintext highlighter-rouge">map</code> and <code class="language-plaintext highlighter-rouge">filter</code> example, which is very
short, very readable, and even a bit… <a href="https://en.wikipedia.org/wiki/Functional_programming">functional</a>.<sup style="anchor-name:--fnref-reduce" id="fnref:reduce"><a href="#fn:reduce" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:reduce">
      <p>But what about <code class="language-plaintext highlighter-rouge">reduce()</code>, the third function of the classic “filter-map-reduce” triplet? Python does have a reduce function, but it was moved to <code class="language-plaintext highlighter-rouge">functools.reduce()</code> because <a href="https://www.python.org/dev/peps/pep-3100/#built-in-namespace">“a loop is more readable most of the time”</a>. <a href="#fnref:reduce" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[For loops are great, but I am a big fan of replacing them with simple functions. Python provides a couple of building blocks.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/patterns/naturalists_misc_vol_1_painted_snake.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/patterns/naturalists_misc_vol_1_painted_snake.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Jupyter Notebook Templates for Data Science: Plotting</title>
      <link href="https://alexgude.com/blog/data-science-plotting-notebook-template/" rel="alternate" type="text/html" title="Jupyter Notebook Templates for Data Science: Plotting" />
      <published>2020-07-27T00:00:00-07:00</published>
      <updated>2020-07-27T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_plotting_notebook_template</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-plotting-notebook-template/"><![CDATA[<p>I recently released my <a href="https://github.com/agude/Jupyter-Notebook-Template-Library">Jupyter Notebook Template Library</a>. Its goal
is to accelerate your data science projects without having to spend hours
poring over old notebooks to find handy code snippets. In this post I dive
into the plotting notebook to show you what it can do.</p>

<h2 id="the-plotting-notebook">The Plotting Notebook</h2>

<p>Visualizing your data is a critical step in understanding it, and so it is
appropriate that the <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/d6cda39c388154cb8f4073e669efff109c743a99/notebooks/basic-plotting-template.ipynb"><strong>first notebook in the library</strong></a> helps
with making beautiful plots.</p>

<p>The notebook begins with boilerplate code that defines metadata for the
resulting files and also changes some defaults, such as the figure size and
resolution, font size, and legend frame. After that there are a few helpful
functions which I will discuss below.</p>

<h3 id="draw-bands">Draw Bands</h3>

<p>One of my favorite functions is <code class="language-plaintext highlighter-rouge">draw_bands()</code>. It draws a set of alternating colored
bands on the background of the plot based on the axis tick locations.</p>

<p>When called with just the axis, like <code class="language-plaintext highlighter-rouge">draw_bands(ax)</code>, it produces this:</p>

<p><a href="/files/jupyter-library//bands.svg"><img src="/files/jupyter-library//bands.svg" alt="A plot showing the default grey bands." /></a></p>

<p>But you can also customize the color using <code class="language-plaintext highlighter-rouge">draw_bands(ax, color="orange",
alpha=0.05)</code>, which produces:</p>

<p><a href="/files/jupyter-library//orange_bands.svg"><img src="/files/jupyter-library//orange_bands.svg" alt="A plot showing the orange bands." /></a></p>

<p>These bands are a subtle way of indicating where on the X-axis a point lies,
which is especially useful when plotting a time series. I use them often. Here
are some examples:</p>

<ul>
  <li>
    <p><a href="/blog/my-sons-language-development-comparison/#development"><strong>Discussing my sons’ language development</strong></a> to highlight each month.</p>
  </li>
  <li>
    <p><a href="/blog/hour-record-plot-improvements/#improvements"><strong>Plotting the progression of the cycling hour record</strong></a> to show each decade.</p>
  </li>
  <li>
    <p><a href="/blog/switrs-bicycle-crashes-by-date/#day-by-day"><strong>Exploring when cyclists are involved in traffic accidents</strong></a> to highlight the seasonality.</p>
  </li>
</ul>

<h3 id="draw-legends">Draw Legends</h3>

<p>I like minimal, but informative, legends. Color alone is often enough to
differentiate lines or points, so I wrote a function to change the color of
the legend text to match the line, called <code class="language-plaintext highlighter-rouge">draw_colored_legend()</code>. It produces
a legend like on this plot:</p>

<p><a href="/files/jupyter-library//legend.svg"><img src="/files/jupyter-library//legend.svg" alt="A plot showing my colored legend." /></a></p>

<p>This legend style can be seen in these posts:</p>

<ul>
  <li>
    <p><a href="/blog/my-second-sons-words/#the-words"><strong>Plotting my son’s language development</strong></a> to label each language.</p>
  </li>
  <li>
    <p><a href="/blog/tdf-prize-money-plot-improvements/#improvements"><strong>Plotting Tour de France Prize Money</strong></a> to label the winner’s prize compared to the total.</p>
  </li>
  <li>
    <p><a href="/blog/data-science-salaries-by-gender/#by-region"><strong>Comparing Data Science Salaries by Gender</strong></a> to differentiate the points for men and women.</p>
  </li>
</ul>

<h2 id="putting-it-together">Putting It Together</h2>

<p>The <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/d6cda39c388154cb8f4073e669efff109c743a99/notebooks/basic-plotting-template.ipynb">plotting notebook</a> enables you to make beautiful plots
quickly and easily. For example, this plot:</p>

<p><a href="/files/jupyter-library//example_plot.svg"><img src="/files/jupyter-library//example_plot.svg" alt="An example plot from the notebook library" /></a></p>

<p>Was produced by this short code snippet:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">fig</span><span class="p">,</span> <span class="n">ax</span> <span class="o">=</span> <span class="nf">setup_plot</span><span class="p">(</span>
    <span class="n">title</span><span class="o">=</span><span class="sh">"</span><span class="s">Title</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">xlabel</span><span class="o">=</span><span class="sh">"</span><span class="s">X-axis</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">ylabel</span><span class="o">=</span><span class="sh">"</span><span class="s">Y-axis</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="n">ax</span><span class="p">.</span><span class="nf">scatter</span><span class="p">(</span><span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">)</span><span class="o">-</span><span class="mf">0.65</span><span class="p">,</span> <span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">),</span> <span class="n">label</span><span class="o">=</span><span class="sh">"</span><span class="s">First dataset</span><span class="sh">"</span><span class="p">)</span>
<span class="n">ax</span><span class="p">.</span><span class="nf">scatter</span><span class="p">(</span><span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">)</span><span class="o">-</span><span class="mf">0.35</span><span class="p">,</span> <span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">),</span> <span class="n">label</span><span class="o">=</span><span class="sh">"</span><span class="s">Second dataset</span><span class="sh">"</span><span class="p">)</span>

<span class="nf">draw_colored_legend</span><span class="p">(</span><span class="n">ax</span><span class="p">)</span>

<span class="nf">draw_bands</span><span class="p">(</span><span class="n">ax</span><span class="p">)</span>

<span class="nf">save_plot</span><span class="p">(</span><span class="n">fig</span><span class="p">,</span> <span class="sh">"</span><span class="s">/tmp/output.svg</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>If the notebook template library is useful to you, be sure to let me know on
<a href="https://twitter.com/alex_gude/">Twitter</a> or <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/issues">Github</a>. Your feedback helps make the project
better for everyone!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
          <category term="data-science" />
        
          <category term="jupyter" />
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[Jumpstart your visualizations with this Jupyter plotting notebook!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/jupyter-library/jupiter_red_spot_juno.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/jupyter-library/jupiter_red_spot_juno.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Machine Learning Deployment: Shadow Mode</title>
      <link href="https://alexgude.com/blog/machine-learning-deployment-shadow-mode/" rel="alternate" type="text/html" title="Machine Learning Deployment: Shadow Mode" />
      <published>2020-06-30T00:00:00-07:00</published>
      <updated>2020-06-30T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/machine_learning_deployment_shadow_mode</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/machine-learning-deployment-shadow-mode/"><![CDATA[<p>Deploying a machine learning product so that it can be used is essential to
getting value out of it. But it is one of the hardest parts of building the
product.</p>

<p>In this post I will focus on a small piece of deployment: <em>“How do I test my
new model in production?”</em> One answer, and a method I often employ when
initially deploying models, is <strong>shadow mode</strong>.</p>

<p>If you’re interested in a broader overview of building and deploying machine
learning products, I highly recommend <a href="https://mlpowered.com/">Emmanuel Ameisen’s</a> book:
<a href="https://mlpowered.com/book/"><em>Building Machine Learning Powered Applications</em></a>!<sup style="anchor-name:--fnref-disclaimer" id="fnref:disclaimer"><a href="#fn:disclaimer" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<h2 id="what-is-shadow-mode">What Is Shadow Mode?</h2>

<p>To launch a model in shadow mode, you deploy the new, shadow model alongside
the old, live model.<sup style="anchor-name:--fnref-live" id="fnref:live"><a href="#fn:live" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> The live model continues to handle all requests,
but the shadow model also runs on some (or all) of the requests. This allows
you to safely test the new model against real data while avoiding the risk of
service disruptions.</p>

<h2 id="when-would-i-use-shadow-mode">When Would I Use Shadow Mode?</h2>

<p>Shadow mode is a great way to test a few things:</p>

<ul>
  <li>
    <p><strong>Engineering</strong>: With a shadow model you can test that the “pipeline” is
working: the model is getting the inputs it expects, and it is returning
results in the correct format. You can also verify that the latency is not too
high.</p>
  </li>
  <li>
    <p><strong>Outputs</strong>: You can verify that the distribution of results looks the way
you expect (for example, your model is not reporting just a single value for
all input).</p>
  </li>
  <li>
    <p><strong>Performance</strong>: You can verify that the shadow model is producing results
that are comparable to or better than those of the live model. This involves
choosing the right <a href="/blog/machine-learning-metrics-interview/">machine learning metric based on the business use
case</a> to compare the model’s outputs against ground truth or
against each other.</p>
  </li>
</ul>

<p>Shadow mode works well when the result of the model does not need a user
action to validate it. Models where you try to influence the user—for
example a recommendation model where success means more sales converted—are
better tested using an <a href="https://en.wikipedia.org/wiki/A/B_testing">A/B test</a>. The big difference between an A/B test
and shadow mode is that in an A/B test traffic is split between the two models
whereas in shadow mode the two models operate on the same events.</p>

<h2 id="how-do-i-deploy-in-shadow-mode">How Do I Deploy In Shadow Mode?</h2>

<p>There are two general methods that I use for deploying in shadow mode. Both
are relative to the <a href="https://en.wikipedia.org/wiki/Application_programming_interface">API</a> for the live model: either <a href="#in-front-of-the-api"><em>in front of the
live API</em></a> or <a href="#behind-the-api"><em>behind the live API</em></a>.</p>

<h3 id="in-front-of-the-api">In Front of the API</h3>

<p>To put a model in shadow mode <em>in front of the API</em>, you host two API
endpoints: one for the live model and one for the shadow model. The caller
makes a call to both of them whenever they would normally call the live model.
The caller can disregard the response, but they should log it so that the
results can be compared. I have drawn this structure below:</p>

<p><a href="/files/shadow-mode//shadow_mode_in_front_of_the_api.svg"><img src="/files/shadow-mode//shadow_mode_in_front_of_the_api.svg" alt="A diagram showing how in front of the API shadow mode is constructed." /></a></p>

<p>This way of deploying is well-suited to situations where the calling team is
change-adverse or has very strict requirements for how the shadow model must
perform because it gives them control. I have found it useful for deploying
models that have a large effect on some <a href="https://en.wikipedia.org/wiki/Conversion_funnel">conversion funnel</a>, like a
model that runs at new user creation and blocks suspected bad actors.</p>

<p>The advantages of this method are:</p>

<ul>
  <li>
    <p><strong>The caller has control.</strong> They decide when to switch the shadow model to
live. They can roll back instantly if there are problems. They can even stop
the experiment if it is hurting their system.</p>
  </li>
  <li>
    <p><strong>The call can be different.</strong> If the shadow model requires different inputs
(perhaps a new ID associated with the user), its API can be different than
that of the live model.</p>
  </li>
</ul>

<p>The main disadvantages are:</p>

<ul>
  <li>
    <p><strong>The change is closer to the customer.</strong> The calling code is generally
closer to the core business, so any bug introduced during integration of the
shadow model is likely to be more impactful.</p>
  </li>
  <li>
    <p><strong>Tighter coordination is required.</strong> The team that owns the model and the
team that calls it will both have to make changes to their code: the model
team to spin up an endpoint, and the calling team to add the call to the
second model as well as a logging action.</p>
  </li>
</ul>

<h3 id="behind-the-api">Behind the API</h3>

<p>To put a model in shadow mode <em>behind the API</em>, you change the code that
responds to API requests to call the live and shadow model. You log the
results of both models<sup style="anchor-name:--fnref-logging" id="fnref:logging"><a href="#fn:logging" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> but only return the result from the live
model. I have drawn a schematic of this below:</p>

<p><a href="/files/shadow-mode//shadow_mode_behind_the_api.svg"><img src="/files/shadow-mode//shadow_mode_behind_the_api.svg" alt="A diagram showing how behind the API shadow mode is constructed." /></a></p>

<p>This method is great when you want to move quickly (and break things), because
you can change the shadow model without having to coordinate with the calling
team. To the outside world the API looks unchanged and so hides the testing
going on behind it.</p>

<p>The advantages of this method are:</p>

<ul>
  <li>
    <p><strong>The model host has control.</strong> You can change the shadow model, turn it on,
turn it off, and swap in a new one at a whim. You can log exactly what you are
interested in recording.</p>
  </li>
  <li>
    <p><strong>Little coordination with other teams is required.</strong> To the outside world
the API looks the same as before; no one else has to change their code.</p>
  </li>
</ul>

<p>The main disadvantages are:</p>

<ul>
  <li>
    <p><strong>The shadow model must be input compatible with the live model.</strong> Since the
outside world is not changing what it passes to the API, your shadow model is
restricted to the same inputs as the live model (although it <em>can</em> choose to
use only a subset, or get additional inputs via some other method).</p>
  </li>
  <li>
    <p><strong>You still have to change the calling code.</strong> Eventually, when the model is
ready to replace the live model, you will need to change the API version and
change the calling code to use this new version. This means there is a little
extra work to be done once you are satisfied with your test results.</p>
  </li>
</ul>

<h2 id="conclusion">Conclusion</h2>

<p>Deploying a model in shadow mode is an easy way to test your model on live
data. It is flexible and allows you to empower the right team to control the
experiment.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:disclaimer">
      <p><strong>Disclaimer</strong>: I was a technical editor for the book, but make no money off sales. <a href="#fnref:disclaimer" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:live">

      <p>By <em>“live model”</em>, I mean whatever system is currently doing the job that
the shadow model will do. It could be a model, a heuristic, a simple <code class="language-plaintext highlighter-rouge">if</code>
statement, or even nothing at all. <a href="#fnref:live" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:logging">
      <p>You <em>are</em> logging your live results, right? <a href="#fnref:logging" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="machine-learning" />
        
          <category term="machine-learning-engineering" />
        
      

      

      
      
        <summary type="html"><![CDATA[Deploying machine learning models is hard; Shadow Mode is one way to make testing a little easier.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/shadow-mode/bricks_at_mit.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/shadow-mode/bricks_at_mit.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">A Review of Nookdesk’s Standing Desk</title>
      <link href="https://alexgude.com/blog/nookdesk-review/" rel="alternate" type="text/html" title="A Review of Nookdesk’s Standing Desk" />
      <published>2020-05-27T00:00:00-07:00</published>
      <updated>2020-05-27T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/nookdesk_review</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/nookdesk-review/"><![CDATA[<p>I bought a standing desk recently because I am working from home three days a
week. I considered many desks but I decided on a <a href="https://www.nookdesk.com/">Nookdesk</a>. It was
not easy to choose because I only found <a href="https://macsources.com/nookdesk-review-ordering-and-building-of-the-smart-desk-that-enhances-your-life/">one <strong>real</strong> review
online</a>; all the others were rewritten press releases.</p>

<p>Luckily I have a good friend who had just bought one. She was happy with her
Nookdesk and so I bought mine on the strength of her recommendation. As it turns
out, that was a good decision. For a complete review, read on.</p>

<h2 id="desktop-and-frame">Desktop and Frame</h2>

<p>The <a href="https://www.nookdesk.com/">Nookdesk</a> has a lot of customization options. The desktops are
offered in multiple materials from various laminates to solid bamboo. The
desktop is always 30″ deep with widths from 48″ to 72″. It is 3/4″ thick. The
desk frame comes in three colors: white, black, and grey.</p>

<p>I chose a 30″ by 72″ (+$200) blackwood laminate (+$100) desktop with white
frame. The wood laminate feels great: smooth with just a bit of depth from
the grain. The white frame contrasts nicely with the black desktop.</p>

<p><img src="/files/nookdesk//nookdesk_wood_laminate_surface.jpg" alt="The wood laminate surface." /></p>

<p>The desk height ranges from 23.5″ to 49″ as measured from the ground to the
top of the desktop. This is just enough range for me as use it at 26.5″ when
seated and 47.5″ when standing.</p>

<h2 id="storage">Storage</h2>

<p>The Nookdesk has some unique storage options. One of them is the “upper
storage” add-on which I would describe more as a shelf. It is mounted on the
top of the desk at the back edge. It is 12″ deep and rises 4″ above the
desktop. I store my game controller and headphones under the shelf, and I put
my monitors on top to bring them up to eye-level. The upper storage add-on
cost $100, but was well worth the expense.</p>

<p><img src="/files/nookdesk//nookdesk_upper_storage.jpg" alt="The upper storage add-on to the nookdesk." /></p>

<p>There is a similar “lower storage” shelf that mounts under the desk at the
front edge. I did not order one because I like to have my desk surface very
close to my knees in the sitting position, and the lower storage would add 4″
to the bottom of the desk. Like the “upper storage”, it also costs $100.</p>

<p>The ability to add some storage to my desk and raise my monitors was the
feature that originally drew me to the Nookdesk and was what finally sold me
on it.</p>

<h2 id="accessories">Accessories</h2>

<p>Nookdesk offers a wide variety of accessories, from keyboard trays to
speakers. Below I review only the ones I purchased.</p>

<h3 id="controller">Controller</h3>

<p><img src="/files/nookdesk//nookdesk_controller.jpg" alt="The desk motor controller." /></p>

<p>The Nookdesk comes with one of two controllers: a simple one which has only up
and down buttons, and a programmable one with three memory settings for an
additional $34. I got the programmable controller because getting my desk to
exactly the right height each time was important to me. I only really needed
two memory positions: one for standing and one for sitting. I still have not
assigned the third position.</p>

<h3 id="powerstrip">Powerstrip</h3>

<p><img src="/files/nookdesk//nookdesk_freeport_powerstrip.jpg" alt="The freeport powerstrip mounted under the desk." /></p>

<p>Nookdesk sells a Freeport power strip that attaches to the underside of the
desk for $80. It has eight outlets (including two with space around them for
those extra-large plugs) and four USB-A outlets, three of which provide up
2.4A and one which supports <a href="https://en.wikipedia.org/wiki/Quick_Charge">Quick Charge 3.0</a>. It has a 10 foot cord.</p>

<p>You can plug the desk’s motor into one of the eight outlets so that you only
have to plug in the Freeport to the wall. I am satisfied with the power strip,
but I suspect I could have found a cheaper one from a third party.</p>

<h3 id="cpu-mount">CPU Mount</h3>

<p>Nookdesk also sells an under-desk CPU mount. It fits towers up to 9.25″ wide
and 30″ tall, although if the desk is in its lowest position there is only 20″
of clearance for the tower.</p>

<p><img src="/files/nookdesk//nookdesk_cpu_mount.jpg" alt="The under-desk CPU mount." /></p>

<p>I love the CPU mount; it is one of my favorite features of the desk. It was
well worth the $100 it cost as it keeps my computer off the ground and frees
up space on my desktop. I recommend it highly!</p>

<h3 id="cable-tray">Cable Tray</h3>

<p><img src="/files/nookdesk//nookdesk_cable_tray.jpg" alt="The under-desk cable tray." /></p>

<p>The cable management tray is a metal tray that runs along the bottom of the
desk at the rear edge and allows you to tuck your cables into it to keep them
tidy. It costs $100. It does a good job of keeping cables from hanging down
and getting in your way when you’re sitting, but seems expensive. However, I
don’t know how I would organize my cables without it.</p>

<h2 id="ordering-and-shipping">Ordering and Shipping</h2>

<p>You can order the desk without creating an account on Nookdesk’s website, but
if you do not create an account then retrieving your order status is
difficult. You are supposed to use a special page on the site where you enter
your email address, but it did not work for me. Instead I had to get customer
support to send me the tracking number.</p>

<p>Delivery was free (due to a <a href="https://www.evodesk.com/deals">coupon code, <code class="language-plaintext highlighter-rouge">NOOK150</code></a>) but seems to run
about $120 normally. The desk was sent by FedEx in two shipments totaling six
boxes. The two shipments arrived on the same day a week after I ordered.</p>

<p>The box containing the desktop was very large and relatively heavy. I was able
to lug it up my stairs by myself but it would have been much easier to have a
second person. The box containing the desk frame was also really heavy
(perhaps more so) because of the two motors that move the desk up and down,
but it was thankfully pretty compact and so easy to maneuver.</p>

<h2 id="assembly">Assembly</h2>

<p>Assembling the desk was actually harder than I expected because there were a
few problems:</p>

<ul>
  <li>
    <p>The hardware—screws, bolts, guides, and whatnot—was packed in unlabeled
bags and the instructions reused the same icons to indicate different screws.</p>
  </li>
  <li>
    <p>The bill of materials listed both items that came in the bags and those that
came preinstalled in the same count. When it said “10 screws of type X” it
might mean there were 4 in the bag and 6 already attached to the frame.</p>
  </li>
  <li>
    <p>Hardware was included that is only used for optional accessories. I had
screws leftover that I had not used at all. This made me worry that I had
completely forgotten some critical step.</p>
  </li>
  <li>
    <p>Two hex keys are included, but one of them had such poor tolerance that it
would not drive the bolts. I ended up using a set of hex keys that I already
owned.</p>
  </li>
  <li>
    <p>The shelf is connected to the frame with four screws, but the contact area
is so small it kept coming apart. I needed to add washers that I had lying
around to make a solid connection. See my photo below:</p>
  </li>
</ul>

<p><img src="/files/nookdesk//nookdesk_shelf_washer.jpg" alt="Attaching the upper storage with washers." /></p>

<p>Assembly took between 3 to 4 hours, but was not particularly difficult after I
worked through the above issues. I made a minor mistake where I did not
install the frame rails at the right time and had to backtrack and remove some
pieces to get them in. <a href="https://macsources.com/nookdesk-review-ordering-and-building-of-the-smart-desk-that-enhances-your-life/">Jon Walters</a> says he made the same mistake in
his review.</p>

<p>You need enough space to flip the desk over during assembly because the build
is started with the desk upside down. The desk is unwieldy, so it helps to
have a second person for this step, but I was able to do the entire build
solo.</p>

<h2 id="final-thoughts">Final Thoughts</h2>

<p>With taxes and shipping my desk was $1057, after a coupon code gave me $150
off and free shipping. Expensive, but I spend 10 hours a day at it so the
price seems fair.</p>

<p>Although the assembly and shipping experience could be improved, the desk is
solid and includes great accessories. I have really appreciated being able to
switch positions instead of sitting for 8 hours, and so I am very happy with
my Nookdesk.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[I bought a Nookdesk standing desk now that I'm working from home; I like it! Read on for a detailed review.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/nookdesk/nookdesk.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/nookdesk/nookdesk.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Jupyter Notebook Templates for Data Science</title>
      <link href="https://alexgude.com/blog/data-science-notebook-templates/" rel="alternate" type="text/html" title="Jupyter Notebook Templates for Data Science" />
      <published>2020-04-27T00:00:00-07:00</published>
      <updated>2020-04-27T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_notebook_templates</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-notebook-templates/"><![CDATA[<p>I love Jupyter notebooks (even if <a href="/blog/jupyter-not-for-development/">I have strong opinions about their
misuse</a>) and so I use them constantly, both at work and here in my
articles. They are the best way to explore a dataset and make visualizations.</p>

<p>But my workflow with notebooks is not very efficient; it is made up of the
following steps:</p>

<ol>
  <li>
    <p>Start a brand new, <em>completely empty</em> notebook.</p>
  </li>
  <li>
    <p>Load the data and start cleaning it.</p>
  </li>
  <li>
    <p>Begin making plots.</p>
  </li>
  <li>
    <p>Realize I already have some code from a different project to make nice plots.</p>
  </li>
  <li>
    <p>Dig through my repositories looking for the code.</p>
  </li>
  <li>
    <p>Copy and paste the first code I find that sort of does what I need (and
which probably is not the most recent or nicest version).</p>
  </li>
  <li>
    <p>Hack the code up and make it even uglier.</p>
  </li>
</ol>

<p>After five years, I am ready for something better. And so I present the <a href="https://github.com/agude/Jupyter-Notebook-Template-Library"><strong>Jupyter
Notebook Template Library</strong></a>, which has revolutionized my workflow and which
I gladly share with you as well, gentle reader.</p>

<h2 id="jupyter-notebook-template-library">Jupyter Notebook Template Library</h2>

<p>The <a href="https://github.com/agude/Jupyter-Notebook-Template-Library">Jupyter Notebook Template Library</a> is a repository of notebook
templates, each targeted at a different use case. The templates let me get
right to working with the data as quickly as possible. And the library
guarantees that my notebook will always have the latest and greatest helper
functions without having to dig through my old work.</p>

<h3 id="the-plotting-template">The Plotting Template</h3>

<p>The first notebook in the library is the <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/8c13dc10c4dbcf724357857692ab7ac64fb83e09/notebooks/basic-plotting-template.ipynb"><strong>Plotting Template</strong></a>.
Its goal is to change the above workflow to this:</p>

<ol>
  <li>
    <p>Download the right template.</p>
  </li>
  <li>
    <p>Load data and start cleaning it.</p>
  </li>
  <li>
    <p>Make <strong>beautiful</strong> plots.</p>
  </li>
</ol>

<p>It lets me write this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">fig</span><span class="p">,</span> <span class="n">ax</span> <span class="o">=</span> <span class="nf">setup_plot</span><span class="p">(</span>
    <span class="n">title</span><span class="o">=</span><span class="sh">"</span><span class="s">Title</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">xlabel</span><span class="o">=</span><span class="sh">"</span><span class="s">X-axis</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">ylabel</span><span class="o">=</span><span class="sh">"</span><span class="s">Y-axis</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="n">ax</span><span class="p">.</span><span class="nf">scatter</span><span class="p">(</span><span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">)</span><span class="o">-</span><span class="mf">0.65</span><span class="p">,</span> <span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">),</span> <span class="n">label</span><span class="o">=</span><span class="sh">"</span><span class="s">First dataset</span><span class="sh">"</span><span class="p">)</span>
<span class="n">ax</span><span class="p">.</span><span class="nf">scatter</span><span class="p">(</span><span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">)</span><span class="o">-</span><span class="mf">0.35</span><span class="p">,</span> <span class="n">np</span><span class="p">.</span><span class="n">random</span><span class="p">.</span><span class="nf">rand</span><span class="p">(</span><span class="mi">500</span><span class="p">),</span> <span class="n">label</span><span class="o">=</span><span class="sh">"</span><span class="s">Second dataset</span><span class="sh">"</span><span class="p">)</span>

<span class="nf">draw_colored_legend</span><span class="p">(</span><span class="n">ax</span><span class="p">)</span>

<span class="nf">draw_bands</span><span class="p">(</span><span class="n">ax</span><span class="p">)</span>

<span class="nf">save_plot</span><span class="p">(</span><span class="n">fig</span><span class="p">,</span> <span class="sh">"</span><span class="s">/tmp/output.svg</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>And get this, complete with curated font sizes, a patented striped background,
and a focused legend:</p>

<p><a href="/files/jupyter-library//example_plot.svg"><img src="/files/jupyter-library//example_plot.svg" alt="An example plot from the notebook library" /></a></p>

<p>You can read more about the plotting notebook in detail here:</p>

<!-- A grid of hand-selected related posts. -->

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/data-science-plotting-notebook-template/">
      <img src="https://alexgude.com/files/jupyter-library/jupiter_red_spot_juno.jpg" alt="The planet Jupiter as seen by the Juno spacecraft.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/data-science-plotting-notebook-template/">
      <strong>Jupyter Notebook Templates for Data Science: Plotting</strong>
    </a>
<br />
Jumpstart your visualizations with this Jupyter plotting notebook!  </div>
</li>
</ul>

<h3 id="the-time-series-plotting-template">The Time Series Plotting Template</h3>

<p>The <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/blob/master/notebooks/basic-time-series-plotting-template.ipynb"><strong>second notebook in the library</strong></a> helps you take a
dataframe of events and turn it into a time series plot with each item broken
out into its own line.</p>

<p>You just write:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="n">seaborn</span> <span class="k">as</span> <span class="n">sns</span>

<span class="n">fig</span><span class="p">,</span> <span class="n">ax</span> <span class="o">=</span> <span class="nf">setup_plot</span><span class="p">(</span><span class="n">title</span><span class="o">=</span><span class="sh">"</span><span class="s">Collisions by Make</span><span class="sh">"</span><span class="p">)</span>

<span class="n">pivot</span> <span class="o">=</span> <span class="nf">plot_time_series</span><span class="p">(</span><span class="n">df</span><span class="p">,</span> <span class="n">ax</span><span class="p">,</span> <span class="n">date_col</span><span class="o">=</span><span class="n">DATE_COL</span><span class="p">,</span> <span class="n">category_col</span><span class="o">=</span><span class="sh">"</span><span class="s">vehicle_make</span><span class="sh">"</span><span class="p">,</span> <span class="n">resample_frequency</span><span class="o">=</span><span class="sh">"</span><span class="s">W</span><span class="sh">"</span><span class="p">)</span>

<span class="c1"># Move labels slightly to avoid overlap
</span><span class="n">nudges</span> <span class="o">=</span> <span class="p">{</span><span class="sh">"</span><span class="s">Toyota</span><span class="sh">"</span><span class="p">:</span> <span class="mi">15</span><span class="p">,</span> <span class="sh">"</span><span class="s">Honda</span><span class="sh">"</span><span class="p">:</span> <span class="o">-</span><span class="mi">8</span><span class="p">}</span>
<span class="nf">draw_left_legend</span><span class="p">(</span><span class="n">ax</span><span class="p">,</span> <span class="n">nudges</span><span class="o">=</span><span class="n">nudges</span><span class="p">,</span> <span class="n">fontsize</span><span class="o">=</span><span class="mi">25</span><span class="p">)</span>

<span class="n">sns</span><span class="p">.</span><span class="nf">despine</span><span class="p">(</span><span class="n">trim</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>

<span class="nf">save_plot</span><span class="p">(</span><span class="n">fig</span><span class="p">,</span> <span class="sh">"</span><span class="s">/tmp/output.svg</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>To make this plot:</p>

<p><a href="/files/jupyter-library//make_collision_in_time.svg"><img src="/files/jupyter-library//make_collision_in_time.svg" alt="An example plot from the time series notebook library" /></a></p>

<p>You can read more about it here:</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/data-science-timeseries-plotting-notebook-template/">
      <img src="https://alexgude.com/files/jupyter-library/jupiter_in_the_rearview_mirror.jpg" alt="The planet Jupiter as seen by the departing Juno spacecraft.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/data-science-timeseries-plotting-notebook-template/">
      <strong>Jupyter Notebook Templates for Data Science: Plotting Time Series</strong>
    </a>
<br />
Jumpstart your time series visualizations with this Jupyter plotting notebook!  </div>
</li>
</ul>

<h2 id="conclusion">Conclusion</h2>

<p>Enjoy the templates, I hope they make you more productive! And if you are
feeling generous, I would love <a href="https://github.com/agude/Jupyter-Notebook-Template-Library/issues">contributions</a>!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="jupyter" />
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[Jupyter notebooks are great for data exploration; jumpstart your work with this library of useful notebook templates!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/jupyter-library/jupiter_cassini_20001229.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/jupyter-library/jupiter_cassini_20001229.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Terribly Clever(ly Terrible) Code</title>
      <link href="https://alexgude.com/blog/cleverly-worst-code/" rel="alternate" type="text/html" title="My Terribly Clever(ly Terrible) Code" />
      <published>2020-03-30T00:00:00-07:00</published>
      <updated>2020-03-30T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/cleverly_worst_code</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/cleverly-worst-code/"><![CDATA[<p>I started learning C++ in graduate school. I had written Python for five
years, so I thought I was pretty good at writing code and thinking through
problems.<sup style="anchor-name:--fnref-wasnt" id="fnref:wasnt"><a href="#fn:wasnt" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Like many new programmers I enjoyed finding clever solutions
to problems; sometimes too clever. This is the story of one of those times.</p>

<h2 id="the-problem">The Problem</h2>

<p>I had to handle setting some state based on two variables. Each variable could
take one of a few discrete values. For simplicity, think of the variables and
possible values:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Variable Name</th>
      <th style="text-align: right">Possible Values</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">direction</code></td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">north</code>, <code class="language-plaintext highlighter-rouge">east</code>, <code class="language-plaintext highlighter-rouge">south</code>, <code class="language-plaintext highlighter-rouge">west</code></td>
    </tr>
    <tr>
      <td style="text-align: left"><code class="language-plaintext highlighter-rouge">travel_mode</code></td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">bike</code>, <code class="language-plaintext highlighter-rouge">car</code>, <code class="language-plaintext highlighter-rouge">plane</code></td>
    </tr>
  </tbody>
</table>

<p>All twelve combinations required doing something slightly different, so the
first code I wrote looked something like this:</p>

<div class="language-cpp highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">if</span> <span class="p">(</span><span class="n">direction</span> <span class="o">==</span> <span class="s">"north"</span> <span class="o">&amp;&amp;</span> <span class="n">travel_mode</span> <span class="o">==</span> <span class="s">"bike"</span><span class="p">)</span> <span class="p">{</span>
  <span class="n">do_north_bike_stuff</span><span class="p">();</span>
<span class="p">}</span>
<span class="k">else</span> <span class="nf">if</span> <span class="p">(</span><span class="n">direction</span> <span class="o">==</span> <span class="s">"north"</span> <span class="o">&amp;&amp;</span> <span class="n">travel_mode</span> <span class="o">==</span> <span class="s">"car"</span><span class="p">)</span> <span class="p">{</span>
  <span class="n">do_north_car_stuff</span><span class="p">();</span>
<span class="p">}</span>
<span class="k">else</span> <span class="nf">if</span> <span class="p">(</span> <span class="p">...</span> <span class="p">)</span> <span class="p">{</span>
  <span class="p">...</span> <span class="c1">// etc.</span>
<span class="p">}</span>
</code></pre></div></div>

<p>This code wasn’t clever; it was boring and repetitive so I looked for a way to
rewrite it! I had recently learned about a cool way to replace <a href="https://en.cppreference.com/w/cpp/language/if"><code class="language-plaintext highlighter-rouge">if/else</code></a>
in C++: the <a href="https://en.cppreference.com/w/cpp/language/switch"><code class="language-plaintext highlighter-rouge">switch</code></a> statement. I had to use it!</p>

<h2 id="making-it-worse">Making It Worse</h2>

<p>But a <code class="language-plaintext highlighter-rouge">switch</code> statement needs integral values, so I had to map each state to
numbers. Easy enough, but I quickly ran into a problem: I had to switch based
on both values, so I had to combine the integers in some manner. Then I
remembered a “useful” math fact: the <a href="https://en.wikipedia.org/wiki/Fundamental_theorem_of_arithmetic">product of unique primes is itself
unique</a>.<sup style="anchor-name:--fnref-useful" id="fnref:useful"><a href="#fn:useful" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> A horrible plan came together, it looked like this:</p>

<div class="language-cpp highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kt">int</span> <span class="n">NORTH</span> <span class="o">=</span> <span class="mi">2</span><span class="p">;</span>
<span class="kt">int</span> <span class="n">EAST</span> <span class="o">=</span> <span class="mi">3</span><span class="p">;</span>
<span class="kt">int</span> <span class="n">SOUTH</span> <span class="o">=</span> <span class="mi">5</span><span class="p">;</span>
<span class="kt">int</span> <span class="n">WEST</span> <span class="o">=</span> <span class="mi">7</span><span class="o">:</span>

<span class="kt">int</span> <span class="n">BIKE</span> <span class="o">=</span> <span class="mi">11</span><span class="p">;</span>
<span class="kt">int</span> <span class="n">CAR</span> <span class="o">=</span> <span class="mi">13</span><span class="p">;</span>
<span class="kt">int</span> <span class="n">PLANE</span> <span class="o">=</span> <span class="mi">17</span><span class="p">;</span>

<span class="k">switch</span><span class="p">(</span><span class="n">direction</span> <span class="o">*</span> <span class="n">travel_mode</span><span class="p">)</span> <span class="p">{</span>
  <span class="k">case</span> <span class="n">NORTH</span> <span class="o">*</span> <span class="n">BIKE</span><span class="p">:</span>  <span class="n">do_north_bike_stuff</span><span class="p">();</span> <span class="k">break</span><span class="p">;</span>
  <span class="k">case</span> <span class="n">EAST</span>  <span class="o">*</span> <span class="n">BIKE</span><span class="p">:</span>  <span class="n">do_east_bike_stuff</span><span class="p">();</span> <span class="k">break</span><span class="p">;</span>
  <span class="p">...</span> <span class="c1">// etc.</span>
  <span class="k">case</span> <span class="n">WEST</span>  <span class="o">*</span> <span class="n">PLANE</span><span class="p">:</span> <span class="n">do_west_plane_stuff</span><span class="p">();</span> <span class="k">break</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>My code was actually even worse; if you are morbidly curious I <a href="/blog/cleverly-worst-code/the-code-itself/">archived it
here</a>. I did not assign nice readable variables like <code class="language-plaintext highlighter-rouge">NORTH</code> but just
used the numbers, so it looked like <code class="language-plaintext highlighter-rouge">case 2 * 11: do_north_bike_stuff()</code>.</p>

<p>This code is <strong>way</strong> too clever; needing number theory to understand
control flow is a <em>huge</em> warning sign. With ten years more experience, I
actually prefer the verbose but understandable <code class="language-plaintext highlighter-rouge">if/else</code> method.</p>

<h2 id="a-better-way">A Better Way</h2>

<p>So how would I write it now? I think a hybrid method is actually the way to
go, using <a href="https://en.cppreference.com/w/cpp/language/enum"><code class="language-plaintext highlighter-rouge">enum</code></a><sup style="anchor-name:--fnref-post" id="fnref:post"><a href="#fn:post" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> and a few helper functions:</p>

<div class="language-cpp highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">enum</span> <span class="n">TravelMode</span> <span class="p">{</span> <span class="n">BIKE</span><span class="p">,</span> <span class="n">CAR</span><span class="p">,</span> <span class="n">PLANE</span> <span class="p">};</span>
<span class="k">enum</span> <span class="n">Direction</span> <span class="p">{</span> <span class="n">NORTH</span><span class="p">,</span> <span class="n">EAST</span><span class="p">,</span> <span class="n">SOUTH</span><span class="p">,</span> <span class="n">WEST</span> <span class="p">};</span>

<span class="k">switch</span><span class="p">(</span><span class="n">travel_mode</span><span class="p">)</span> <span class="p">{</span>
  <span class="k">case</span> <span class="n">BIKE</span><span class="p">:</span>  <span class="n">do_bike_stuff</span><span class="p">(</span><span class="n">direction</span><span class="p">);</span> <span class="k">break</span><span class="p">;</span>
  <span class="k">case</span> <span class="n">CAR</span><span class="p">:</span>   <span class="n">do_car_stuff</span><span class="p">(</span><span class="n">direction</span><span class="p">);</span> <span class="k">break</span><span class="p">;</span>
  <span class="k">case</span> <span class="n">PLANE</span><span class="p">:</span> <span class="n">do_plane_stuff</span><span class="p">(</span><span class="n">direction</span><span class="p">);</span> <span class="k">break</span><span class="p">;</span>
<span class="p">}</span>

<span class="kt">void</span> <span class="nf">do_bike_stuff</span><span class="p">(</span><span class="n">Direction</span> <span class="n">dir</span><span class="p">)</span> <span class="p">{</span>
  <span class="k">switch</span><span class="p">(</span><span class="n">dir</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">case</span> <span class="n">NORTH</span><span class="p">:</span> <span class="p">...;</span> <span class="k">break</span><span class="p">;</span>
    <span class="k">case</span> <span class="n">EAST</span><span class="p">:</span> <span class="p">...;</span> <span class="k">break</span><span class="p">;</span>
    <span class="p">...</span> <span class="c1">// etc.</span>
  <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>This has a few nice advantages:</p>

<ul>
  <li>
    <p>It delegates the complexity of handling the direction to each travel mode.
This is logical because it’s likely the way a car handles East and West are
very similar, and very different from how a plane would.</p>
  </li>
  <li>
    <p>By using <code class="language-plaintext highlighter-rouge">enum</code> the compiler can check that we handle every case. If we
forget <code class="language-plaintext highlighter-rouge">case BIKE</code> or <code class="language-plaintext highlighter-rouge">case EAST</code>, the compiler can warn us.</p>
  </li>
</ul>

<p>In the end, readable is better than clever, even if you have a bunch more
lines to read!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:wasnt">
      <p><em>Narrator</em>: He wasn’t. <a href="#fnref:wasnt" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:useful">
      <p><em>Narrator</em>: It was not useful. <a href="#fnref:useful" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:post">

      <p>See <a href="/blog/python-patterns-enum/"><em>Python Patterns: Enum</em></a>, which covers use-cases
for <code class="language-plaintext highlighter-rouge">enum</code> in Python. It works essentially the same in C++. <a href="#fnref:post" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[When I was young and naive I tried to write very clever code. Here is one of the worst examples.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/worst-code/montreal_light_head_and_power_consolidated_linesmen_1928.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/worst-code/montreal_light_head_and_power_consolidated_linesmen_1928.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Comparison of My Two Sons’ Language Development</title>
      <link href="https://alexgude.com/blog/my-sons-language-development-comparison/" rel="alternate" type="text/html" title="Comparison of My Two Sons’ Language Development" />
      <published>2020-02-10T00:00:00-08:00</published>
      <updated>2020-02-10T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/my_sons_language_development_comparison</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-sons-language-development-comparison/"><![CDATA[<p>My son Theo was born in the summer of 2016 and my son Cory was born in the
winter of 2017. Our family is multi-lingual so we knew our sons would
therefore have complicated and interesting language development. My wife and I
are (unsurprisingly) <strong>huge nerds</strong> so we wrote down each new word they
learned so we could explore how they learned language. I wrote <a href="/blog/my-sons-words/">a post
focusing on Theo’s language development</a> and <a href="/blog/my-second-sons-words/">another post focusing
on Cory’s language development</a>; this month I compare them.</p>

<h2 id="the-data">The Data</h2>

<p>The data was collected by my wife and I attempting to identify when the boys
had learned a new word and writing it down. The most common error is writing
down words one of the boys does not yet really know. This would lead to an
increase in the number of words known at any time in the data; still the
difference between the boys should be unaffected as the error pushes the data
in the same direction for both of them.</p>

<p>I discuss data collection in more depth in <a href="/blog/my-sons-words/#the-data">Theo’s</a> and
<a href="/blog/my-second-sons-words/#the-data">Cory’s</a> data sections. You can find the Jupyter notebook used
to perform this analysis <a href="/files/my-sons-words-comparison//Theo%20vs%20Cory%20words.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/my-sons-words-comparison//Theo%20vs%20Cory%20words.ipynb">rendered on Github</a>).
The data can be found <a href="/files/my-sons-words/theo_words.csv">here</a> and <a href="/files/my-second-sons-words/cory_words.csv">here</a>.</p>

<h2 id="development">Development</h2>

<p>Below I have plotted the number of words each of my sons knew as a function of
their age. Theo, our first son, is represented by dashed lines; Cory, our
second son, is represented by solid lines.</p>

<p><a href="/files/my-sons-words-comparison//child0_vs_child1_total_words_linear.svg"><img src="/files/my-sons-words-comparison//child0_vs_child1_total_words_linear.svg" alt="A plot showing the number of words my sons could speak as a function of
their age." /></a></p>

<p>Second children are known to have slower onset of language
development.<sup style="anchor-name:--fnref-pine" id="fnref:pine"><a href="#fn:pine" class="footnote" rel="footnote" role="doc-noteref">1</a></sup><sup style="anchor-name:--fnref-berglund" id="fnref:berglund"><a href="#fn:berglund" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> They learn their first 50 known words more
slowly, but catch up to their older siblings quickly, learning their first 100
words at about the same age. The advantage is small though; the average
difference in time to first 50 words between the first and second child is
only 1 month.</p>

<p>Theo and Cory do not follow this trend. Cory was 3 to 4 months faster than
Theo to hit language development milestones in Cantonese, English, and
Spanish; he also knew many more animal sounds. Theo was artificially limited
in his Spanish acquisition though, as <a href="/blog/my-sons-words/">mentioned in his post</a>,
because my mother had injured herself and so Theo could not visit my parents
for a few months.</p>

<p>Sign is the only area where Theo eventually learned faster than Cory. I
suspect this is because he needed sign to communicate as he did not know as
many words as quickly, whereas Cory gave up on sign once he could talk.</p>

<p>Finally, I owe Theo’s doctor an apology: she was always a little worried about
Theo’s language development, but we did not take it seriously because we knew
that bilingual children developed language more slowly. Looking at the data
and comparing to his brother, I think our pediatrician was right to be
worried. Thankfully, Theo has had no problems since and now talks incessantly.</p>

<h2 id="other-writings-on-language-development">Other Writings on Language Development</h2>

<p>If you enjoyed this article, here are all the other articles I wrote about
<a href="/topics/childhood-language/">language development</a>!</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" alt="Black and white photo of two young boys hanging out a window, their faces smudged with soot.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <strong>Comparison of My Three Sons’ Language Development</strong>
    </a>
<br />
I recorded the words my sons spoke as they learned our various languages and now I compare how each developed! Read on to find out how each son learned.  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <img src="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" alt="A woodcut by J. W. Orr showing a father giving his son a picture book in a richly appointed study.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <strong>My Third Son’s Language Development</strong>
    </a>
<br />
We tracked my third son's language development word by word. Here, in plots, is how he learned to speak. Take a look!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <img src="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a man using a blackboard to teach young children punctuation.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <strong>My Second Son’s Language Development</strong>
    </a>
<br />
My second son is a little over two years old. We tracked every word he's spoken to watch his language development, and now you can observe it too!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <img src="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a woman using a blackboard to teach young children how to pronounce words.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <strong>My Son’s Language Development</strong>
    </a>
<br />
My son is a little over two and unfortunately he has two huge nerds for parents. We tracked every word he's spoken to watch his language development, and now you can join us!  </div>
</li>
</ul>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:pine">

      <p><span class="citation">Pine, J. M. <a href="https://doi.org/10.2307/1131205">“Variation in vocabulary development as a function of birth order”</a> <cite>Child Development</cite>. vol. 66, no. 1. 1995. pp. 272–281. doi: <a href="https://doi.org/10.2307/1131205">10.2307/1131205</a>.</span> <a href="#fnref:pine" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:berglund">

      <p><span class="citation">E. Berglund, M. Eriksson, and M. Westerlund. <a href="https://doi.org/10.1111/j.1467-9450.2005.00480.x">“Communicative skills in relation to gender, birth order, childcare and socioeconomic status in 18‐month‐old children”</a> <cite>Scandinavian Journal of Psychology</cite>. vol. 46. 2005. pp. 485–491. doi: <a href="https://doi.org/10.1111/j.1467-9450.2005.00480.x">10.1111/j.1467-9450.2005.00480.x</a>.</span> <a href="#fnref:berglund" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="childhood-language" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[Being a nerd dad, I recorded all the words my first two sons spoke as they learned them. Now, I compare their language development rate!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Second Son’s Language Development</title>
      <link href="https://alexgude.com/blog/my-second-sons-words/" rel="alternate" type="text/html" title="My Second Son’s Language Development" />
      <published>2020-01-30T00:00:00-08:00</published>
      <updated>2020-01-30T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/my_second_sons_words</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-second-sons-words/"><![CDATA[<p>My second son, Cory, was born in the winter of 2017. Like <a href="/blog/my-sons-words/">my first
son</a>, we tracked his language development to see how fast he picked
up the various languages that our family speaks.</p>

<h2 id="the-data">The Data</h2>

<p>The data was collected in the <a href="/blog/my-sons-words/#the-data">same manner as last time</a>. The
only difference was I added a new language category for “animal sounds”
because when recording that data for Theo I realized it was hard to tell “moo”
from “mu”. A catch-all category for animals made data logging much easier.</p>

<p>I had the same data collection difficulties as last time: when Cory was young,
it was hard to decide if he was associating a sound with a concept. When Cory
was older, he was so good at imitating sounds that it was hard to know if he
knew the word or was just repeating what you had said. There was a new
difficulty this time as well: we had our third child just as Cory’s language
development exploded. Cory was with his grandparents for a few weeks as we
adjusted to the new baby and so I was not able to record his new words during
that time period.</p>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/my-second-sons-words//Cory's%20first%20words.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/my-second-sons-words//Cory's%20first%20words.ipynb">rendered on Github</a>). The data can be found
<a href="/files/my-second-sons-words//cory_words.csv">here</a>.</p>

<h2 id="development">Development</h2>

<p>Cory’s first <a href="https://en.wikipedia.org/wiki/Baby_sign_language">baby sign</a> was waving goodbye at 11 months. He
learned four more—pointing at things he wanted, blowing kisses, shaking his
head for “no”, and shaking his hands for “all done”—before speaking his
first word. His first word, <a href="/blog/my-sons-words/#development">like his older brother</a>, was in
Cantonese, although it was “dad” and not “dog”. It was two more months before
he said “mom” in Cantonese, at which point he already knew “older brother”,
“ball”, “there”, and “pick me up”. My wife was disappointed it took so long,
but I suspect it was because he spent so much time with her he did not need a
word to get her attention, he always had it.</p>

<p>Cory’s sign language development outpaced his other languages until he was 20
months old. My suspicion is the same as it was with his brother: sign was his
universal language. English could only be used to talk to dad, Cantonese could
only be used to talk to mom, but sign could communicate with both of us and
even his grandparents.</p>

<p>Cory’s language development is plotted below, showing the number of words he
could speak in each “language” as a function of how old he was.</p>

<p><a href="/files/my-second-sons-words//child1_total_words_linear.svg"><img src="/files/my-second-sons-words//child1_total_words_linear.svg" alt="A plot showing the number of words my second son could speak as a function
of age." /></a></p>

<p>Cantonese and English exploded at about 19 months; Cory doubled his vocabulary
in those two languages in just four weeks. He almost doubled it again in his
20th month. In his 22nd month, he picked up about 50 Cantonese words and
almost 70 English ones!</p>

<p>Part of that growth in the 22nd month is a data entry artifact: that is when
he went to stay with his grandparents. He had a week to learn words and I only
recorded them when he came back to us and spoke them. You can see two flat
areas in his English and Cantonese language development—one right before 22
months and one right after—that are due to this effect.</p>

<p>Spanish started pretty slowly because only my father speaks it to him. At 22
months we moved closer to my parents so Cory spent more time with them. You
can see a bit of an increase in the amount of Spanish learned then; I’m sure
we would see a much larger increase if we kept recording.</p>

<h2 id="the-words">The Words</h2>

<p>I plotted a selection of some of Cory’s first words in each language below.
Notice that I have switched to a log plot for the <em>y</em>-axis to better show the
beginnings of each language.</p>

<p><a href="/files/my-second-sons-words//child1_first_words.svg"><img src="/files/my-second-sons-words//child1_first_words.svg" alt="A plot showing the first words my second son could speak as a function of
age." /></a></p>

<p>A few fun words:</p>

<ul>
  <li>
    <p><strong>Grandpa</strong> (Spanish): As the only person who speaks Spanish full-time to
him, it makes sense that Grandpa would be Cory’s first Spanish word. Cory has
always had a special affinity for my father, so it is appropriate as well.</p>
  </li>
  <li>
    <p><strong>Cow</strong> (English and Spanish) and <strong>Lola</strong> (Spanish): Cory started hearing
<em>La Vaca Lola</em> (Lola the cow) and quickly learned most of the words. He still
loves cows, and it is still one of his favorite songs.</p>
  </li>
  <li>
    <p><strong>Older Brother</strong> (Cantonese): Theo is one of the defining features of
Cory’s life as both Cory’s best friend and primary antagonizer. Cory learned
to identify him quickly (and often would just point and say “big brother” when
Theo had pushed him over).</p>
  </li>
  <li>
    <p><strong>Older Sister</strong> (Cantonese): Cory doesn’t have a sister, but there were two
girls who lived next door to our apartment with whom he would play and call
“older sister”.</p>
  </li>
  <li>
    <p><strong>Lion Dancer</strong> (Cantonese): Theo became obsessed with lion dancers during
Lunar New Year last year. Cory has picked up on this obsession as well and
often does a solo lion dance.</p>
  </li>
  <li>
    <p><strong>Google</strong> and <strong>Deebot</strong> (English): We have a lot of technology in our
house and Cory has learned all about it. We talk to our Google Home devices
several times a day trying to get them to play music or animal sounds, and
Deebot vacuums the kitchen every night as Cory watches in awe.</p>
  </li>
</ul>

<h2 id="other-writings-on-language-development">Other Writings on Language Development</h2>

<p>If you enjoyed this article, here are all the other articles I wrote about
<a href="/topics/childhood-language/">language development</a>!</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" alt="Black and white photo of two young boys hanging out a window, their faces smudged with soot.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <strong>Comparison of My Three Sons’ Language Development</strong>
    </a>
<br />
I recorded the words my sons spoke as they learned our various languages and now I compare how each developed! Read on to find out how each son learned.  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <img src="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" alt="A woodcut by J. W. Orr showing a father giving his son a picture book in a richly appointed study.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <strong>My Third Son’s Language Development</strong>
    </a>
<br />
We tracked my third son's language development word by word. Here, in plots, is how he learned to speak. Take a look!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" alt="Black and white photo of a young boy at a school desk.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <strong>Comparison of My Two Sons’ Language Development</strong>
    </a>
<br />
Being a nerd dad, I recorded all the words my first two sons spoke as they learned them. Now, I compare their language development rate!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <img src="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a woman using a blackboard to teach young children how to pronounce words.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-words/">
      <strong>My Son’s Language Development</strong>
    </a>
<br />
My son is a little over two and unfortunately he has two huge nerds for parents. We tracked every word he's spoken to watch his language development, and now you can join us!  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="childhood-language" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[My second son is a little over two years old. We tracked every word he's spoken to watch his language development, and now you can observe it too!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" />
        <media:content medium="image" url="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Get Things Done by Tracking Them</title>
      <link href="https://alexgude.com/blog/getting-things-done/" rel="alternate" type="text/html" title="Get Things Done by Tracking Them" />
      <published>2019-12-30T00:00:00-08:00</published>
      <updated>2019-12-30T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/getting_things_done</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/getting-things-done/"><![CDATA[<p>As a new data scientist, my life was pretty easy: come in, work on one big
problem, go home. But as I became more comfortable in my role, I started
taking on new responsibilities: mentor new hires, improve internal processes,
attend planning meetings and follow up, and so on.</p>

<p>By the time I was running a team, I no longer had one thing I was focusing on
a day; instead, I had dozens. Often these things would only take 15–30 minutes
to do, but I would have to fit them in with all my meetings. My strategy of
“remember what I need to do today” stopped working because I just had too many
things to do, and items were constantly added and removed.</p>

<p>I needed a strategy; what I came up with was a version of <a href="https://en.wikipedia.org/wiki/David_Allen_(author)">David
Allen’s</a> <a href="https://en.wikipedia.org/wiki/Getting_Things_Done">Getting Things Done</a>. Here is how I do it.</p>

<h2 id="trello">Trello</h2>

<p>The key insight of the Getting Things Done framework is that there should be
<strong>one</strong> place where tasks are recorded. I was storing tasks in my notes—both
the electronic and paper versions—in my email using the “important” flag,
and in my brain which meant I had to check multiple places when I had an
opening in my schedule to complete a task. I decided the one place to track
tasks would be <a href="https://trello.com/">Trello</a> because adding and arranging tasks is easy, it
supports my computer and phone, and has enough <a href="#useful-power-ups">power-ups</a>
that I can use to customize my workflow.</p>

<p>I set up six columns named as follows (yes, with the emoji):</p>

<ul>
  <li>
    <p><strong>🗃️ Backlog</strong>: Tasks which are not urgent or not ready to start; a lot of
“someday” tasks end up here.</p>
  </li>
  <li>
    <p><strong>📥 Inbox</strong>: The landing place for new tasks where I focus on quick entry,
not on correctness.</p>
  </li>
  <li>
    <p><strong>☀️ Do This Week</strong>: Tasks to do this week.</p>
  </li>
  <li>
    <p><strong>📅 Do Today</strong>: Tasks to do today.</p>
  </li>
  <li>
    <p><strong>✅ Done</strong>: Finished tasks.</p>
  </li>
  <li>
    <p><strong>🎓 Learning</strong>: Books, articles, and papers that I want to read, classes I
want to take, etc.</p>
  </li>
</ul>

<h2 id="triage">Triage</h2>

<p>Every morning I triage my tasks by taking one of two actions:</p>

<ol>
  <li>
    <p>If the task can be done in 5 minutes, I do it immediately.</p>
  </li>
  <li>
    <p>If the task can’t be completed quickly, I move it to the correct column and
make sure all the meta data like title are descriptive.</p>
  </li>
</ol>

<p>Triaging mostly means moving tasks from the <strong>Inbox</strong> and the <strong>Backlog</strong> to
one of the <strong>Do</strong> columns. On Monday I also make sure to take a moment to
think about my week and fill up the <strong>Do This Week</strong> column. Everyday I think
about what I’m going to need to get done that day and move those tasks to the
<strong>Do Today</strong> column.</p>

<p>When I have time to complete some tasks, I look at <strong>Do Today</strong> and pick
something I can fit in, and do it. Then I move the task to <strong>Done</strong>.</p>

<p>The <strong>Learning</strong> column is special, it is where I store links to things I want
to read, but did not have time to when I encountered them. If I’m honest, it’s
mostly a place where I store links I find on Twitter when I’m on my phone.</p>

<h2 id="capturing-tasks">Capturing Tasks</h2>

<p>A good friend told me that the key to this system was that <em>“you can’t block
on IO”</em>, meaning it must be easy to record a task and that you should not
worry about getting it recorded perfectly. Triage is when you can clean up any
typos and add context.</p>

<p>There are three places where I capture tasks:</p>

<ul>
  <li>
    <p><strong>My Computer</strong>: Here I often write out a long task and put it directly into
the correct column because the computer makes data entry easy.</p>
  </li>
  <li>
    <p><strong>In My Email</strong>: I use an Outlook plugin that takes an email and turns it
directly into a task in Trello. These tasks end up in the <em>Inbox</em> and I triage
them later. This lets me go through the dozens (or hundreds) of emails and
quickly get the ones that require action turned into a task.</p>
  </li>
  <li>
    <p><strong>On The Go</strong>: I use the Trello app on my phone, and I have a widget that
lets me add a card with the press of a button. Just like email, this puts it
into the <em>Inbox</em> column for later triage. I find using the dictation feature
of my phone makes adding tasks really fast.</p>
  </li>
</ul>

<h2 id="useful-power-ups">Useful Power-Ups</h2>

<p>Trello has some useful addons called “power-ups”. I found two power-ups to be
useful (although the free plan only allows you to use one power-up at a time):</p>

<ul>
  <li>
    <p><strong>Card Repeater</strong>: This power-up lets you schedule a card to reappear. I
used it for things that have to happen on a fixed schedule like “plan for the
weekly team meeting” and “fill out the project status update”. I <em>also</em> used
it for things I wanted to get in the habit of doing, like “ask for feedback”.
To this end I set three “ask for feedback” tasks appear on my board every
Monday so I could get in the habit of asking.</p>
  </li>
  <li>
    <p><strong>Card Aging</strong>: This power-up faded cards that hadn’t been touched for a
while. This let me see what tasks were being ignored and let me re-evaluate
if they were still worth doing.</p>
  </li>
</ul>

<p>This is how I get things done as a data scientist. It is pretty simple, but
effective, and easy to put your own spin on it. Give it a try!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
      

      

      
      
        <summary type="html"><![CDATA[As I've gotten more senior as a data scientist, I've found I have to keep track of more and more things. This is how I do it!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/getting-things-done/sticky-notes-by-irfan-simsar.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/getting-things-done/sticky-notes-by-irfan-simsar.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Improving Wikipedia’s Tour de France Prize Money Plot</title>
      <link href="https://alexgude.com/blog/tdf-prize-money-plot-improvements/" rel="alternate" type="text/html" title="Improving Wikipedia’s Tour de France Prize Money Plot" />
      <published>2019-11-25T00:00:00-08:00</published>
      <updated>2019-11-25T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/tdf_prize_money_plot_improvements</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/tdf-prize-money-plot-improvements/"><![CDATA[<p>The Tour de France is the most important bike race of the year, and it is
therefore the race with the most prize money awarded. Wikipedia has this plot
showing how that prize money has grown over the years:</p>

<figure>
  
  <a href="/files/tdf-prize-money//TdFPrizeMoney.svg">
    <img src="/files/tdf-prize-money//TdFPrizeMoney.svg" alt="A line plot showing the Total and Winner's prize money in the
  Tour de France over its history." decoding="async" />
  </a>
  
  
  
  <figcaption><a href="https://en.wikipedia.org/wiki/File:TdFPrizeMoney.svg"><em>Prize money in the Tour de France</em></a>, ©<a href="https://en.wikipedia.org/wiki/User:EdgeNavidad">EdgeNavidad</a>
  (<a href="https://en.wikipedia.org/wiki/Public_domain">Public Domain</a>)</figcaption>
  
</figure>

<p>The plot is pretty good, at least at first glance! It is (appropriately) a
<a href="https://en.wikipedia.org/wiki/Semi-log_plot">log plot</a>.<sup style="anchor-name:--fnref-exp" id="fnref:exp"><a href="#fn:exp" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> It labels all of its pieces. It even has gaps for
when the race was not held. But the plot also has a few problems:</p>

<ul>
  <li>
    <p>The X-axis is wrong; the race did not start before 1900 and the two gaps are
from the World Wars which did not happen in 1898 and 1921.</p>
  </li>
  <li>
    <p>The text is too small to read easily at Wikipedia’s default 200px image
size.</p>
  </li>
  <li>
    <p>The axis labels are redundant and tick labels have a lot of zeroes.</p>
  </li>
</ul>

<p>I decided to fix it up using <a href="https://www.bikeraceinfo.com/tdf/tdf-prizes.html">data from Bike Race Info</a>, like I
did <a href="/blog/hour-record-plot-improvements/">last time when I fixed the <em>Hour Record Plot</em></a>.</p>

<h2 id="improvements">Improvements</h2>

<p>Here is my version:</p>

<p><a href="/files/tdf-prize-money//tdf_prize_money_in_2013_euro.svg"><img src="/files/tdf-prize-money//tdf_prize_money_in_2013_euro.svg" alt="The same information as above, but using a step plot with better labeling." /></a></p>

<p>The code that generated the improved plots can be found <a href="/files/tdf-prize-money//Tour%20de%20France%20Prize%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/tdf-prize-money//Tour%20de%20France%20Prize%20Plot.ipynb">rendered on Github</a>). The data <a href="/files/tdf-prize-money//tdf_prizes_dataframe.json">is here</a>, and the code
that cleaned the data <a href="/files/tdf-prize-money//bikeraceinfo.com%20Scraper.ipynb">is here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/tdf-prize-money//bikeraceinfo.com%20Scraper.ipynb">rendered on
Github</a>).</p>

<p>I fixed the X-axis so that the dates are now correct! The first race is in
1903 as expected. I have also removed the axis label because I think the tick
labels make it clear what is plotted. I have added my (<a href="/blog/data-science-plotting-notebook-template/#draw-bands">patent pending
😛</a>) grey stripes to the background to indicate each decade.</p>

<p>I changed the Y-axis to be more readable by abbreviating the numbers using
<em>K</em> and <em>M</em>. I also removed the label and replaced it with the euro symbol (€)
on each tick.</p>

<p>I made all the text larger and the lines thicker to improve legibility when the plot
is downscaled. I have also changed from a line plot to a step plot because the amount
of prize money changes at specific moments in time, not continually.</p>

<p>Finally, I have cleaned up the data a bit. The original plot used uncorrected
Euro even though the original prizes were in old Franc, new Franc, and Euro
depending on the year. I have normalized all values to
2013 Euro. I have included this information in the subtitle so that it
survives even if the plot is separated from its caption on Wikipedia.</p>

<p>Overall I think it is an improvement, so I have contributed it back to the
community <a href="https://commons.wikimedia.org/wiki/File:Tdf_prize_money_in_2013_euro.svg">here</a>.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:exp">
      <p>Inflation is exponential. <a href="#fnref:exp" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[Time to improve another plot from Wikipedia. This time I tackle one showing the prize money in the Tour de France over time!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/tdf-prize-money/godinat_and_level_in_1934.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/tdf-prize-money/godinat_and_level_in_1934.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Interview Question: What Machine Learning Metric to Use</title>
      <link href="https://alexgude.com/blog/machine-learning-metrics-interview/" rel="alternate" type="text/html" title="Interview Question: What Machine Learning Metric to Use" />
      <published>2019-10-28T00:00:00-07:00</published>
      <updated>2019-10-28T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/machine_learning_metrics_interview</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/machine-learning-metrics-interview/"><![CDATA[<p>As part of our interview cycle, candidates work with some data and build a
simple model. After we talk through the modeling and data work, I ask them to
come up with a business case for the model. Once they have done so, I follow
up with:</p>

<blockquote>
  <p>How would you measure the success of this model in production?</p>
</blockquote>

<p>I have heard a lot of answers. They generally fall into two categories:
machine learning theory focused, and business focused. The first is good, the
second is better. I will go through each below.</p>

<p>To have something concrete to discuss, we will consider the following problem:
“Train a model to classify customers as <em>‘high value customer’</em>, and use it to
decide if we are going to show them an up-sell page.”</p>

<h2 id="machine-learning-theory-focused">Machine Learning Theory Focused</h2>

<p>A common answer, especially from more junior interviewees, is “Accuracy”,
which is the number of things your model classifies correctly divided by the
total number of things.</p>

<p><strong>Accuracy is rarely a good answer for real world problems</strong>, because the
classes are often imbalanced. If only 1 in 1000 users is a <em>‘high value
customer’</em>, then a model that only returns ‘false’, despite having an accuracy
of 99.9%, would be clearly worthless.</p>

<p>When pushed about accuracy’s obvious shortcomings, the candidate may fall back
to something like F1 score, which is the harmonic mean of precision and
recall. <strong>F1 is a lazy answer</strong>. It is better than accuracy because it is
relatively less sensitive to imbalanced data, but rarely are precision and
recall equally important (how the customer interacts with the model determines
their relative weighting) and so the harmonic mean is not often applicable.</p>

<h2 id="customer-focused">Customer Focused</h2>

<p>The best candidates consider the problem more deeply. Instead of jumping to
find a metric, they start by thinking about the experience of using a model
from a business or user perspective. A good way to frame this is “What does a
false positive cost my user?” and “How does that compare to the cost of a
false negative?”</p>

<p>Sometimes a <strong>false positive is costly</strong>, as might be the case if the
resulting action is drastic, like shutting down a user’s account. To avoid
that, we need to be confident when we take action that we are only targeting
the right users. In such a scenario, <a href="https://en.wikipedia.org/wiki/Precision_and_recall#Precision">precision</a> is more important
since we want to make sure that most of the events the model flags are true
positives. This is not the situation for our <em>‘high value customer’</em> model,
because showing a user an up-sell page is unlikely to hurt them or cause them
to churn.</p>

<p>Instead, in our case a <strong>false negative is costly</strong>, because we lose the
chance at a large revenue increase from the up-sell. <a href="https://en.wikipedia.org/wiki/Precision_and_recall#Recall">Recall</a> is more
important, as we would rather show a few extra users our up-sell page than
miss the chance to convert a sale.</p>

<p>There are many other metrics that might be useful. The important part is not
so much the metric itself, but what is motivating it, which should be the
business use case and customer experience.</p>

<h2 id="a-great-metric-dollars">A Great Metric: Dollars</h2>

<p>A great metric is the formalized version of our customer focused one: dollars.
Assigning a dollar value to each model result (true positive, false positive,
etc.) would allow us to optimize for revenue. This is often doable in simple
models (like our <em>‘high value customer’</em> model), but in more complicated ones
(like the fraud models I work on) it can be difficult.<sup style="anchor-name:--fnref-rep" id="fnref:rep"><a href="#fn:rep" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> For an example of
using dollars in model optimization, see Airbnb’s post on <a href="https://medium.com/airbnb-engineering/fighting-financial-fraud-with-targeted-friction-82d950d8900e"><em>Fighting Financial
Fraud with Targeted Friction</em></a>.</p>

<p>In short, a good metric is one which ties closely to the business or customer
use. Consider their point of view when answering a metrics question.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:rep">
      <p>One of our largest costs is <a href="https://en.wikipedia.org/wiki/Reputational_risk">reputational risk</a>, which is very hard to assign a number to. <a href="#fnref:rep" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="machine-learning" />
        
          <category term="interviewing" />
        
          <category term="interview-prep" />
        
      

      

      
      
        <summary type="html"><![CDATA[One of my favorite questions to ask in an interview is "What metric should you use to decide if your model works?". Read on to find out what a good answer looks like!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/interviews/Artgate_Fondazione_Cariplo_-_Canova_Antonio,_Allegoria_della_Giustizia.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/interviews/Artgate_Fondazione_Cariplo_-_Canova_Antonio,_Allegoria_della_Giustizia.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Using Travis Build Stages to Test Multiple Python Versions and Publish to Pypi</title>
      <link href="https://alexgude.com/blog/python-pypi-and-travis-staging/" rel="alternate" type="text/html" title="Using Travis Build Stages to Test Multiple Python Versions and Publish to Pypi" />
      <published>2019-09-02T00:00:00-07:00</published>
      <updated>2019-09-02T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/python_pypi_and_travis_staging</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-pypi-and-travis-staging/"><![CDATA[<p>I recently finished a new Python program, <a href="https://github.com/agude/wayback-machine-archiver"><code class="language-plaintext highlighter-rouge">wayback-machine-archiver</code></a>.
See my <a href="/blog/wayback-machine-archiver/">recent post for details</a>. It supports <strong>many</strong> different
versions of Python,<sup style="anchor-name:--fnref-1" id="fnref:1"><a href="#fn:1" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> which required me to <a href="https://en.wikipedia.org/wiki/Build_automation">automate the build system</a>.
I needed my build system to:</p>

<ul>
  <li>
    <p>Install the software for each supported Python version.</p>
  </li>
  <li>
    <p>Run tests against each Python version.</p>
  </li>
  <li>
    <p>Publish <strong><em>exactly once</em></strong> to <a href="https://pypi.org/project/wayback-machine-archiver/">Pypi</a> when all the tests in all
the Python versions had passed.</p>
  </li>
</ul>

<p>Running a bunch of tests, and then deploying your software is exactly what
<a href="https://docs.travis-ci.com/user/build-stages/">Travis Stages</a> were designed for. There was just one problem:
I couldn’t find a good example of how to do it with multiple Python versions.
This post will explain how to do it.</p>

<h2 id="the-travis-configuration">The Travis Configuration</h2>

<p>Here is the full <a href="https://github.com/agude/wayback-machine-archiver/blob/b3d0955e03a09662c1eb9ea962e527a8299bc209/.travis.yml">Travis Configuration file</a>. I’ll go through a
simplified version below that covers just the essentials:</p>

<h3 id="run-tests-in-parallel">Run Tests in Parallel</h3>

<p>We start by setting up the Python versions to use for testing:</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">language</span><span class="pi">:</span> <span class="s">python</span>
<span class="na">dist</span><span class="pi">:</span> <span class="s">xenial</span> <span class="c1"># Required for Python &gt;= 3.7</span>
<span class="na">python</span><span class="pi">:</span>
  <span class="pi">-</span> <span class="s2">"</span><span class="s">2.7"</span>
  <span class="pi">-</span> <span class="s2">"</span><span class="s">3.7"</span>
  <span class="c1"># Also test pypy</span>
  <span class="pi">-</span> <span class="s2">"</span><span class="s">pypy3"</span>
</code></pre></div></div>

<p>This will run tests in parallel on Python 2.7, 3.7, and <a href="https://pypy.org/">Pypy3</a>. Python
3.7 is <a href="https://docs.travis-ci.com/user/languages/python/#python-37-and-higher">only supported on Ubuntu Xenial</a>, so we set that as the
<code class="language-plaintext highlighter-rouge">dist</code>. I removed a bunch of versions for clarity; to add more, just write them
in the list.</p>

<p>Next, we tell Travis how to set up the environment and test the code:</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">install</span><span class="pi">:</span>
  <span class="pi">-</span> <span class="s">pip install -r requirements.txt</span>
<span class="na">script</span><span class="pi">:</span>
  <span class="c1"># Unit tests</span>
  <span class="pi">-</span> <span class="s">python -m pytest -v</span>
  <span class="c1"># Install and smoke test</span>
  <span class="pi">-</span> <span class="s">pip install .</span>
  <span class="pi">-</span> <span class="s">archiver --help</span>
</code></pre></div></div>

<p>This installs the dependencies, runs the unit tests, makes
sure we can pip install the package, and finally runs a quick <a href="https://en.wikipedia.org/wiki/Smoke_testing_(software)">“smoke test”</a> on
the installed package.</p>

<h3 id="build-and-deploy">Build and Deploy</h3>

<p>After the tests succeed (and <em>only</em> after) we build the Pypi package:</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">jobs</span><span class="pi">:</span>
  <span class="na">include</span><span class="pi">:</span>
    <span class="pi">-</span> <span class="na">stage</span><span class="pi">:</span> <span class="s">build</span>
      <span class="na">python</span><span class="pi">:</span> <span class="s2">"</span><span class="s">3.7"</span>
      <span class="na">script</span><span class="pi">:</span> <span class="s">echo "Starting Pypi build"</span>
      <span class="na">deploy</span><span class="pi">:</span>
        <span class="na">provider</span><span class="pi">:</span> <span class="s">pypi</span>
        <span class="na">user</span><span class="pi">:</span> <span class="s">alexgude</span>
        <span class="na">password</span><span class="pi">:</span>
          <span class="na">secure</span><span class="pi">:</span> <span class="s">Bq6I8x...sqslR</span> <span class="c1"># Hashed password</span>
        <span class="na">distributions</span><span class="pi">:</span> <span class="s2">"</span><span class="s">sdist</span><span class="nv"> </span><span class="s">bdist_wheel"</span>
        <span class="na">on</span><span class="pi">:</span>
          <span class="na">tags</span><span class="pi">:</span> <span class="kc">true</span>
          <span class="na">branch</span><span class="pi">:</span> <span class="s">master</span>
          <span class="na">repo</span><span class="pi">:</span> <span class="s">agude/wayback-machine-archiver</span>
        <span class="na">skip_existing</span><span class="pi">:</span> <span class="kc">true</span>
</code></pre></div></div>

<p>This defines a new stage to build the package in 3.7 and then deploys it to
Pypi, but only if it is the master branch on a tagged (from <a href="https://git-scm.com/book/en/v2/Git-Basics-Tagging"><code class="language-plaintext highlighter-rouge">git
tag</code></a>) release.</p>

<p>Which gives us this:<sup style="anchor-name:--fnref-note" id="fnref:note"><a href="#fn:note" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>

<p><a href="/files/travis//results.png"><img src="/files/travis//results.png" alt="A screen shot of the resulting Travis run from this
configuration file." /></a></p>

<p>I hope that helps you set up your own Python packages for testing and
deployment! In the future, I hope to migrate to <a href="https://github.com/features/actions">Github
Actions</a>, but that is for another time.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1">
      <p>Currently 2.7, 3.4 through 3.7, the development versions of 3.7 and 3.8, the nightly release, and <a href="https://pypy.org/">pypy 2.7 and 3.5</a>. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:note">
      <p>I took out a bunch of the versions in the example YAML configuration; the screen shot shows all the versions I test against. <a href="#fnref:note" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Often when building packages, we want to test against multiple versions of the language, and then build the package once. I will show you how to accomplish this using Travis Stages.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/travis/us_airforce_construction_by_sue_sapp.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/travis/us_airforce_construction_by_sue_sapp.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Plotting the 2019 Tour de France</title>
      <link href="https://alexgude.com/blog/2019-tour-de-france-plot/" rel="alternate" type="text/html" title="Plotting the 2019 Tour de France" />
      <published>2019-08-05T00:00:00-07:00</published>
      <updated>2019-08-05T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/2019_tour_de_france_plot</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/2019-tour-de-france-plot/"><![CDATA[<p>There is no bigger event in cycling than the <a href="https://en.wikipedia.org/wiki/Tour_de_France">Tour de France</a>, a race
which takes most of July as it crisscrosses France before bringing the riders
to a fateful final sprint in Paris on the Champs-Élysées. I love both cycling
and plots (as I <a href="/blog/hour-record-plot-improvements/">mentioned last month</a>), so once again I found a
way to combine the two, by graphically exploring how the race unfolded.</p>

<p>The code that generated the plots can be found <a href="/files/tour-de-france//Tour%20de%20France%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/tour-de-france//Tour%20de%20France%20Plot.ipynb">rendered on Github</a>). The data <a href="/files/tour-de-france//2019-tdf-dataframe.json">is here</a>.</p>

<h2 id="the-race-for-yellow">The Race for Yellow</h2>

<p>The <a href="https://en.wikipedia.org/wiki/General_classification_in_the_Tour_de_France">yellow jersey</a> is awarded to the rider with the lowest combined
time across all 21 stages of the tour. Only a few riders are really in
contention for yellow; the vast majority of the others are brought along to
support their team leaders. Going into the 2019 Tour, defending champion
<a href="https://en.wikipedia.org/wiki/Geraint_Thomas">Geraint Thomas</a> was the favorite, but there were several strong
challengers.</p>

<p>Below I show how the top riders did throughout the race by plotting how far
behind the leader they were after each stage. Where a rider’s line is near the
top they are close to taking the lead; when they drop down they are losing
time.</p>

<p><a href="/files/tour-de-france//2019_tour_de_france_top_5.svg"><img src="/files/tour-de-france//2019_tour_de_france_top_5.svg" alt="A line plot showing how far behind the leader each top-finishing rider was after each stage." /></a></p>

<p><a href="https://en.wikipedia.org/wiki/Julian_Alaphilippe">Julian Alaphilippe</a> held the jersey for the most days, even
defending it on the <a href="https://en.wikipedia.org/wiki/Individual_time_trial">individual time trial</a> against expert time trialist
Thomas. But Alaphilippe is not a climber, and after a brave defense of the
jersey in the Pyrenees, he lost time in the Alps to <a href="https://en.wikipedia.org/wiki/Egan_Bernal">Egan Bernal</a>,
<a href="https://en.wikipedia.org/wiki/Steven_Kruijswijk">Steven Kruijswijk</a>, <a href="https://en.wikipedia.org/wiki/Emanuel_Buchmann">Emanuel Buchmann</a>, and <a href="https://en.wikipedia.org/wiki/Thibaut_Pinot">Thibaut
Pinot</a>. Pinot had looked out of contention earlier, but stormed back
with a massive attack on stage 15. Unfortunately, he was forced to withdraw on
stage 19 due to an injury.</p>

<p>Alaphilippe finally fell behind during that stage as well, losing the yellow
jersey to Bernal. He lost his podium spot on stage 20 when he cracked during
the final part of the climb. Alaphilippe finished 5th when the peloton rolled
through the Champs-Élysées.</p>

<h2 id="the-rest-of-the-race">The Rest of the Race</h2>

<p>From the above plot you might think that all the riders in the Tour finish
within a few minutes of each other. But they do not. The last place rider, the
<a href="https://en.wikipedia.org/wiki/Lanterne_rouge">lanterne rouge</a>, was four and a half <strong>hours</strong> behind Egan Bernal.</p>

<p><a href="/files/tour-de-france//2019_tour_de_france.svg"><img src="/files/tour-de-france//2019_tour_de_france.svg" alt="A line plot showing how far behind the leader every rider was for each stage." /></a></p>

<p>Bernal and Alaphilippe, who looked so far apart in the first plot, are now
seen to be neck-and-neck. <a href="https://en.wikipedia.org/wiki/Yoann_Offredo">Yoann Offredo</a> looked like a lock to win
the lanterne rouge when he fell ill on stage 8, but <a href="https://en.wikipedia.org/wiki/Sebastian_Langeveld">Sebastian
Langeveld</a> took it in the penultimate stage after suffering an
injury in the second week of the race.</p>

<p><a href="https://en.wikipedia.org/wiki/Peter_Sagan">Peter Sagan</a>, the <a href="https://en.wikipedia.org/wiki/Points_classification_in_the_Tour_de_France">green jersey</a> winner, was only interested in
sprints. He took it easy on the mountain stages to conserve energy, as you can
see in the steep declines on stages 14 and 15 (in the Pyrenees) and Stages 18–20
(in the Alps). Most other riders performed similarly, although some recovered
time on stage 17 when the favorites let a large breakaway group escape and
gain a 20 minute advantage. Sagan did not make that group, which is clear from
his lack of rise on the plot.</p>

<p><a href="https://en.wikipedia.org/wiki/Romain_Bardet">Romain Bardet</a>, the <a href="https://en.wikipedia.org/wiki/Mountains_classification_in_the_Tour_de_France">polka dot jersey</a> winner, was
fighting for the yellow jersey until stage 14 where he cracked and lost 20
minutes. This forced him to pivot to trying to win the climbing jersey, which
meant he needed to be one of the first riders to reach the top of the
remaining climbs. For the last few climbing stages he stayed with the
favorites or attacked early, keeping his time behind pretty consistent.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[The Tour de France is a sporting event decided by mere minutes; to see exactly how those minutes were earned, read on for my plots!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/tour-de-france/tour_de_france_1932.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Improving Wikipedia’s Hour Record Plot</title>
      <link href="https://alexgude.com/blog/hour-record-plot-improvements/" rel="alternate" type="text/html" title="Improving Wikipedia’s Hour Record Plot" />
      <published>2019-07-09T00:00:00-07:00</published>
      <updated>2019-07-09T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/hour_record_plot_improvements</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/hour-record-plot-improvements/"><![CDATA[<p>The <a href="https://en.wikipedia.org/wiki/Hour_record">cycling hour record</a> is a grueling experience: the would-be
record setter rides as far as they can in one hour. The record was first set
in the modern era by the great <a href="https://en.wikipedia.org/wiki/Eddy_Merckx">Eddy Merckx</a>,<sup style="anchor-name:--fnref-goat" id="fnref:goat"><a href="#fn:goat" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and it has traded
hands multiple times since. Wikipedia has this plot showing the progression:</p>

<figure>
  
  <a href="/files/hour-record//Progression_of_Hour_record_from_Merckx_to_Unified.png">
    <img src="/files/hour-record//Progression_of_Hour_record_from_Merckx_to_Unified.png" alt="A dot plot showing the time and distance for various men's hour records." decoding="async" />
  </a>
  
  
  
  <figcaption><a href="https://en.wikipedia.org/wiki/File:Progression_of_Hour_record_from_Merckx_to_Unified.png"><em>Progression
  of Hour record from Merckx to Unified</em></a>, ©<a href="https://en.wikipedia.org/wiki/User:XyZAn">XyZAn</a> (<a href="https://creativecommons.org/licenses/by-sa/3.0/deed.en">CC-BY-SA
  3.0</a>)</figcaption>
  
</figure>

<p>This plot gets the message across—twice, the distance went up quickly in a short
amount of time—but could be much more effective. Here are some
problems:</p>

<ul>
  <li>
    <p>It is missing a legend and title, which are both necessary to understand it.<sup style="anchor-name:--fnref-plot" id="fnref:plot"><a href="#fn:plot" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>
  </li>
  <li>
    <p>It has too much precision in the date labels, which are down to the day but do
not align with when the records were set.</p>
  </li>
  <li>
    <p>The label text is too small to read easily.</p>
  </li>
  <li>
    <p>It has a lot of unused space.</p>
  </li>
</ul>

<p>I love cycling, and I love plots, so I tried my hand at improving the plot.</p>

<h2 id="improvements">Improvements</h2>

<p>Here is my version:</p>

<p><a href="/files/hour-record//mens_hour_records_progression.svg"><img src="/files/hour-record//mens_hour_records_progression.svg" alt="The same information as above, but using a step plot with better labeling." /></a></p>

<p>I also made the <a href="/files/hour-record//womens_hour_records_progression.svg">same type of plot for the Women’s Hour Record progression</a>.</p>

<p>The code that generated the improved plots can be found <a href="/files/hour-record//Hour%20Record%20Replot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/hour-record//Hour%20Record%20Replot.ipynb">rendered on Github</a>). The data <a href="/files/hour-record//hour_record_dataframe.json">is here</a>.</p>

<p>I added a title and legend. The title makes the subject clear: the progression
of the men’s hour record. The legend is pretty minimal, but conveys that there
are three different types of record, and they are each plotted in a different
color.</p>

<p>The tick labels are now much larger and easier to read. I have changed the
date ticks to every decade because we do not really care about an exact date,
just a rough time and the ordering. I have added a light shading for each
decade to make them easier to tell apart. I have also removed the x-axis label
because it is clear that it shows “years”.</p>

<p>Using the extra white space, I have added <strong>a lot</strong> more information to the
plot: I have added the name of the rider who set each record, and the distance
they rode. I have also added a line indicating the status of each record at
each point in time, making it easy to see where the record is at any point,
and helping to highlight the instances when the record stood for a long time.</p>

<p>This plot took a lot of work to make—matplotlib is not the most forgiving
library—but I think it was worth it. Of course, as a good <a href="https://en.wikipedia.org/wiki/Wikipedia:WikiFairy">WikiFairy</a>, I
<a href="https://en.wikipedia.org/w/index.php?title=Hour_record&amp;oldid=903869466#Statistics">contributed the plots back to Wikipedia</a> so that everyone can
benefit from the improvements!</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:goat">
      <p>The 🐐! <a href="#fnref:goat" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:plot">
      <p>Plots do not need a title or axis labels if the subject is clear without them. In this case though, you would never figure out it was the “Men’s Hour Record Progression” unless someone told you. <a href="#fnref:plot" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="cycling" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[I love Wikipedia, I love cycling, and I love data! So today, I improve Wikipedia's Hour Record Plot! Come take a look!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/hour-record/bicycle_race_by_calvert_litho_co_1895.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/hour-record/bicycle_race_by_calvert_litho_co_1895.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Wayback Machine Archiver: Backup Pages with Python</title>
      <link href="https://alexgude.com/blog/wayback-machine-archiver/" rel="alternate" type="text/html" title="Wayback Machine Archiver: Backup Pages with Python" />
      <published>2019-06-04T00:00:00-07:00</published>
      <updated>2019-06-04T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/wayback_machine_archiver</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/wayback-machine-archiver/"><![CDATA[<p>The <a href="https://archive.org/">Internet Archive</a> runs a service called the <a href="https://archive.org/web/">Wayback Machine</a> to
create a digital archive of the entire internet. It contains the <a href="https://web.archive.org/web/20130518151312/http://alexgude.com/">first
snapshot of my website</a>, as well as many others. I love the ability to
go back and see how my site has evolved, and know that even if I stop hosting
it, the archive will live on. That’s why I want the Wayback Machine to archive
every change my site goes through.</p>

<p>The Internet Archive makes it easy to submit a site for archival, but doing
this manually is time consuming. So I built <a href="https://github.com/agude/wayback-machine-archiver"><strong>Wayback Machine
Archiver</strong></a> to automate it.</p>

<h2 id="wayback-machine-archiver">Wayback Machine Archiver</h2>

<p>The Archiver runs on either <a href="https://www.python.org/">Python</a> 2.7 or 3.4+. It can be installed
with pip:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>wayback-machine-archiver
</code></pre></div></div>

<p>After that, submitting pages for archival is as easy as:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>archiver https://alexgude.com https://alexgude.com/blog/  <span class="c"># etc.</span>
</code></pre></div></div>

<p>This is not much of an improvement over doing it manually, since we still have
to find each URL by hand. Luckily, most sites already have a list of all their
pages in a <a href="https://en.wikipedia.org/wiki/Sitemaps">sitemap.xml</a>.</p>

<p>For example, my sitemap is at <code class="language-plaintext highlighter-rouge">https://alexgude.com/sitemap.xml</code>. It is
automatically generated by <a href="https://en.wikipedia.org/wiki/Jekyll_(software)">Jekyll</a> when I update my site. Archiver
can read a sitemap and submit all of the pages listed in it, making archival
of my blog as easy as:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>archiver <span class="nt">--sitemaps</span> https://alexgude.com/sitemap.xml
</code></pre></div></div>

<p>Archiver has additional options like using multiple threads, logging, and
archiving the sitemap as well as the pages they link to. Checkout the
<a href="https://github.com/agude/wayback-machine-archiver">README</a> on the Github site, or use the <code class="language-plaintext highlighter-rouge">--help</code> flag.</p>

<h3 id="scheduling-archiver">Scheduling Archiver</h3>

<p>Archiver follows the <a href="https://en.wikipedia.org/wiki/Unix_philosophy">Unix Philosophy</a> of “Make each program do one thing
well”, so it leaves the scheduling to another program. I recommend
<a href="https://en.wikipedia.org/wiki/Cron"><code class="language-plaintext highlighter-rouge">cron</code></a>. Here is how I backup my site, and a few others, each day:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code># Backup three websites weekly
@weekly archiver --sitemaps https://alexgude.com/sitemap.xml --archive-sitemap-also --log INFO
@weekly archiver --sitemaps http://charles.uno/sitemap.xml --archive-sitemap-also --log INFO
@weekly archiver https://www.radiokeysmusic.com --sitemaps https://www.radiokeysmusic.com/sitemap.xml --archive-sitemap-also --log INFO
</code></pre></div></div>

<p>If you find my Archiver useful, please consider <a href="https://archive.org/donate/"><strong>donating to the Internet
Archive</strong></a>; none of this would be possible without them! Your company
may even match your donation like mine does!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[The Internet Archive's Wayback Machine tries to keep a complete copy of the internet. With this script, you can submit pages for effortless indexing.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/wayback-machine-archiver/library_of_congress_1902_crop.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/wayback-machine-archiver/library_of_congress_1902_crop.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">The Gender Pay Gap in Data Science Salaries</title>
      <link href="https://alexgude.com/blog/data-science-salaries-by-gender/" rel="alternate" type="text/html" title="The Gender Pay Gap in Data Science Salaries" />
      <published>2019-05-09T00:00:00-07:00</published>
      <updated>2019-05-09T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_salaries_by_gender</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-salaries-by-gender/"><![CDATA[<p>The <a href="https://en.wikipedia.org/wiki/Gender_pay_gap">gender pay gap</a> is a contentious issue, especially in tech where
<a href="https://qz.com/work/1287881/how-technology-companies-alienate-women-during-recruitment/">women are historically excluded</a>. We can explore the gap in
Data Science salaries a little with the same <a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight data</a> I used
<a href="/blog/data-science-salaries/">last time to look at Data Science salaries in general</a>.</p>

<p>Others have looked into the same question before: <a href="https://flolytic.com/">Florian
Lindstaedt</a> used a much larger (but less clean) dataset from Kaggle
to <a href="https://web.archive.org/web/20190825075155/https://flolytic.com/blog/gender-pay-gap-among-data-scientists-on-kaggle">look at the issue on his blog</a>. He found that for data
scientists younger than 30, women earned slightly more, but in the 30–35 age
group men earned more.</p>

<p>My data is much smaller, but better curated. However, it has some biases in
that it is collected from Insight alumni who are mostly:</p>

<ul>
  <li>
    <p>PhDs</p>
  </li>
  <li>
    <p>Early career</p>
  </li>
  <li>
    <p>In high-demand markets</p>
  </li>
  <li>
    <p>Coached on salary negotiation</p>
  </li>
</ul>

<p>Asking the respondent’s gender was added to the survey late, so around a third
of the data does not have that information. This leaves us 79 men and 28
women. Not a huge sample, but better than nothing.</p>

<p>Of course, this low number of women might itself be a further bias: Insight
generally has pretty gender-balanced cohorts, so the fact that many fewer
women have filled out the survey is worrying. It is possible that non-response
is correlated to the underlying distribution, for example, perhaps people who
are paid less refuse to report.</p>

<p>The data used in this post is available <a href="/files/data-science-salaries//insight_salary_survey_cleaned.csv">here</a>. The notebook with all
the code is <a href="/files/data-science-salaries//Data%20Science%20Salary%20Data%20Gender.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/data-science-salaries//Data%20Science%20Salary%20Data%20Gender.ipynb">rendered on Github</a>).</p>

<h2 id="pay-men-vs-women">Pay: Men Vs. Women</h2>

<p>Here is total recurring compensation<sup style="anchor-name:--fnref-salary" id="fnref:salary"><a href="#fn:salary" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> by gender. I have removed all
non-data scientists (like the <a href="/blog/data-science-salaries/#scientists-engineers-and-analysts">MLEs I looked at last time</a>)
because there are very few responses from them. I have also removed the one
data scientist who responded “transgender” without further indicating their
gender identity.</p>

<p>So, how is pay equality in data science?</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_gender.svg"><img src="/files/data-science-salaries//data_science_total_comp_gender.svg" alt="A swarm plot showing salaries for male and female data
scientists." /></a></p>

<p>Pretty equal, actually! The median woman in the sample earns more than the
median man, but of course the number of samples is really small.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Gender</th>
      <th style="text-align: right">Median Total Compensation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Female</td>
      <td style="text-align: right">$149k</td>
    </tr>
    <tr>
      <td style="text-align: left">Male</td>
      <td style="text-align: right">$139k</td>
    </tr>
  </tbody>
</table>

<p>There are lots of things I would like to explore—like “do women see the same benefit from
seniority as men?”, as <a href="/blog/data-science-salaries/#experience-counts-a-lot">I observed last time</a>—but
I just do not have enough women in the sample to say anything conclusive.</p>

<p>Instead I will look at salaries by region (which I know drives large pay
differences) and age, which <a href="https://web.archive.org/web/20190825075155/https://flolytic.com/blog/gender-pay-gap-among-data-scientists-on-kaggle">Florian looked at</a>.</p>

<h2 id="by-region">By Region</h2>

<p>Only California (LA, San Francisco, and Silicon Valley) and the Northeast (New
York, Boston, and DC) have enough respondents to form any reasonable
conclusions, so I limit my sample to those regions.</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_gender_and_location.svg"><img src="/files/data-science-salaries//data_science_total_comp_gender_and_location.svg" alt="A swarm plot showing salaries for male and female data scientists in
California and the East Coast." /></a></p>

<p>Again, these look pretty equal, with the median woman earning slightly more than the
median man in both regions.</p>

<table>
  <thead>
    <tr> <th>Region</th> <th>Gender</th> <th style="text-align: right">Median Total Compensation</th> </tr>
  </thead>
  <tbody>
    <tr> <td rowspan="2">California</td>  <td>Female</td>  <td style="text-align: right">$168k</td> </tr>
    <tr>                                  <td>Male</td>    <td style="text-align: right">$162k</td> </tr>
    <tr> <td rowspan="2">Northeast</td>   <td>Female</td>  <td style="text-align: right">$145k</td> </tr>
    <tr>                                  <td>Male</td>    <td style="text-align: right">$136k</td> </tr>
  </tbody>
</table>

<h2 id="by-age">By Age</h2>

<p>Finally, I can check what Florian found: that women under 30 earned more than
men in the same age range, but men out earned women in the 30–35 age range. I use the same
selection as above, but now partitioning by age instead of region.</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_gender_and_age.svg"><img src="/files/data-science-salaries//data_science_total_comp_gender_and_age.svg" alt="A swarm plot showing salaries for male and female data scientists in
California and the East Coast by age" /></a></p>

<p>I do not see Florian’s trend; instead the salaries look roughly equal, with
the median woman earning more in every age group, as shown below:</p>

<table>
  <thead>
    <tr><th>Age</th> <th>Gender</th> <th style="text-align: right">Median Total Compensation</th></tr>
  </thead>
  <tbody>
    <tr> <td rowspan="2">0 to 30</td>  <td>Female</td>  <td style="text-align: right">$155k</td> </tr>
    <tr>                               <td>Male</td>    <td style="text-align: right">$140k</td> </tr>
    <tr> <td rowspan="2">31–35</td>    <td>Female</td>  <td style="text-align: right">$164k</td> </tr>
    <tr>                               <td>Male</td>    <td style="text-align: right">$148k</td> </tr>
    <tr> <td rowspan="2">36+</td>      <td>Female</td>  <td style="text-align: right">$180k</td> </tr>
    <tr>                               <td>Male</td>    <td style="text-align: right">$138k</td> </tr>
  </tbody>
</table>

<h2 id="conclusion">Conclusion</h2>

<p>In my small dataset, women in data science earn the same as men, and they do
so across regions and age groups. I wish I could have explored more slices of
the data to look at things like seniority, percent of compensation in stock,
etc., but slicing the data very quickly reduces the number of data points
beyond usefulness.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:salary">
      <p>Salary, yearly bonus, and yearly stock grant. Signing bonus is not included. <a href="#fnref:salary" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[How do the salaries of women data scientists compare to those of men? This month we explore pay by gender and location.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-salaries/josef_wagner_hohenberg_the_notary_2_coins.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-salaries/josef_wagner_hohenberg_the_notary_2_coins.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Python Patterns: @total_ordering</title>
      <link href="https://alexgude.com/blog/python-patterns-total-ordering/" rel="alternate" type="text/html" title="Python Patterns: @total_ordering" />
      <published>2019-04-15T00:00:00-07:00</published>
      <updated>2019-04-15T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/python_patterns_total_ordering</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-patterns-total-ordering/"><![CDATA[<p>Python classes come with a set of rich comparison operators. I can compare
strings lexically like so:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="sh">"</span><span class="s">alex</span><span class="sh">"</span> <span class="o">&gt;</span> <span class="sh">"</span><span class="s">alan</span><span class="sh">"</span>
<span class="sh">"</span><span class="s">cat</span><span class="sh">"</span> <span class="o">&lt;</span> <span class="sh">"</span><span class="s">dog</span><span class="sh">"</span>
</code></pre></div></div>

<p>And I can sort numbers including integers and floats:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nf">sorted</span><span class="p">((</span><span class="mi">4</span><span class="p">,</span> <span class="mi">3</span><span class="p">,</span> <span class="mf">2.2</span><span class="p">,</span> <span class="mi">5</span><span class="p">))</span> <span class="o">==</span> <span class="p">[</span><span class="mf">2.2</span><span class="p">,</span> <span class="mi">3</span><span class="p">,</span> <span class="mi">4</span><span class="p">,</span> <span class="mi">5</span><span class="p">]</span>
</code></pre></div></div>

<p>All of these are made possible by <a href="https://docs.python.org/3/reference/datamodel.html#specialnames">special methods</a> defined by each
class. Implementing comparison and sorting for your own classes means defining
six methods, one each for <code class="language-plaintext highlighter-rouge">==</code>, <code class="language-plaintext highlighter-rouge">!=</code>, <code class="language-plaintext highlighter-rouge">&gt;</code>, <code class="language-plaintext highlighter-rouge">&gt;=</code>, <code class="language-plaintext highlighter-rouge">&lt;</code>, and <code class="language-plaintext highlighter-rouge">&lt;=</code>. Thankfully,
Python has a helper method that makes it even simpler: <a href="https://docs.python.org/3/library/functools.html#functools.total_ordering">the <code class="language-plaintext highlighter-rouge">@total_ordering</code>
decorator</a>.</p>

<h2 id="your-library">Your Library</h2>

<p>Let’s make a class to hold books so we can keep track of our library. A basic
<code class="language-plaintext highlighter-rouge">Book</code> class might look like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Book</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">title</span><span class="p">,</span> <span class="n">author</span><span class="p">):</span>
    <span class="n">self</span><span class="p">.</span><span class="n">title</span> <span class="o">=</span> <span class="n">title</span>
    <span class="n">self</span><span class="p">.</span><span class="n">author</span> <span class="o">=</span> <span class="n">author</span>
</code></pre></div></div>

<p>We want the <code class="language-plaintext highlighter-rouge">Book</code> class to be comparable because that will allow us to order
the books on the shelf (using <code class="language-plaintext highlighter-rouge">sorted()</code> for instance). Books will be sorted
first by author and then by title. To implement that, we might write the six special
methods like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Book</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">title</span><span class="p">,</span> <span class="n">author</span><span class="p">):</span>
    <span class="n">self</span><span class="p">.</span><span class="n">title</span> <span class="o">=</span> <span class="n">title</span>
    <span class="n">self</span><span class="p">.</span><span class="n">author</span> <span class="o">=</span> <span class="n">author</span>

  <span class="c1"># Define ==
</span>  <span class="k">def</span> <span class="nf">__eq__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">==</span> <span class="n">theirs</span>

  <span class="c1"># Define !=
</span>  <span class="k">def</span> <span class="nf">__ne__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">!=</span> <span class="n">theirs</span>

  <span class="c1"># Define &lt;
</span>  <span class="k">def</span> <span class="nf">__lt__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">&lt;</span> <span class="n">theirs</span>

  <span class="c1"># Define &lt;=
</span>  <span class="k">def</span> <span class="nf">__le__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">&lt;=</span> <span class="n">theirs</span>

  <span class="c1"># Define &gt;
</span>  <span class="k">def</span> <span class="nf">__gt__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">&gt;</span> <span class="n">theirs</span>

  <span class="c1"># Define &gt;=
</span>  <span class="k">def</span> <span class="nf">__ge__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">&gt;=</span> <span class="n">theirs</span>
</code></pre></div></div>

<p>That is a lot of boilerplate code!</p>

<p>Math tells us that if <code class="language-plaintext highlighter-rouge">self &gt; other</code> is true, then <code class="language-plaintext highlighter-rouge">self &lt; other</code> and <code class="language-plaintext highlighter-rouge">self ==
other</code> are false. We could write our own logic taking advantage of this fact
to reduce the boilerplate, but that is exactly what <code class="language-plaintext highlighter-rouge">@total_ordering</code> from
<code class="language-plaintext highlighter-rouge">functools</code> does already!</p>

<h2 id="with-total_ordering">With <code class="language-plaintext highlighter-rouge">@total_ordering</code></h2>

<p>Using the <code class="language-plaintext highlighter-rouge">@total_ordering</code> decorator<sup style="anchor-name:--fnref-to" id="fnref:to"><a href="#fn:to" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> we only have to define <code class="language-plaintext highlighter-rouge">__eq__</code> and
one of the other comparison methods. The rest of the methods are filled in for
us. It’s used like so:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">functools</span> <span class="kn">import</span> <span class="n">total_ordering</span>


<span class="nd">@total_ordering</span>
<span class="k">class</span> <span class="nc">Book</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">title</span><span class="p">,</span> <span class="n">author</span><span class="p">):</span>
    <span class="n">self</span><span class="p">.</span><span class="n">title</span> <span class="o">=</span> <span class="n">title</span>
    <span class="n">self</span><span class="p">.</span><span class="n">author</span> <span class="o">=</span> <span class="n">author</span>

  <span class="c1"># Define ==
</span>  <span class="k">def</span> <span class="nf">__eq__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">==</span> <span class="n">theirs</span>

  <span class="c1"># Define &lt;
</span>  <span class="k">def</span> <span class="nf">__lt__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="n">ours</span> <span class="o">=</span> <span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">self</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="n">theirs</span> <span class="o">=</span> <span class="p">(</span><span class="n">other</span><span class="p">.</span><span class="n">author</span><span class="p">,</span> <span class="n">other</span><span class="p">.</span><span class="n">title</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">ours</span> <span class="o">&lt;</span> <span class="n">theirs</span>
</code></pre></div></div>

<p>Now we have a much more compact class, but with the same functionality as
before! We can sort our books easily:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">my_books</span> <span class="o">=</span> <span class="p">[</span>
  <span class="nc">Book</span><span class="p">(</span><span class="sh">"</span><span class="s">Absalom, Absalom!</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">William Faulkner</span><span class="sh">"</span><span class="p">),</span>
  <span class="nc">Book</span><span class="p">(</span><span class="sh">"</span><span class="s">The Sun Also Rises</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">Ernest Hemingway</span><span class="sh">"</span><span class="p">),</span>
  <span class="nc">Book</span><span class="p">(</span><span class="sh">"</span><span class="s">For Whom The Bell Tolls</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">Ernest Hemingway</span><span class="sh">"</span><span class="p">),</span>
  <span class="nc">Book</span><span class="p">(</span><span class="sh">"</span><span class="s">The Sound and the Fury</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">William Faulkner</span><span class="sh">"</span><span class="p">),</span>
<span class="p">]</span>

<span class="k">for</span> <span class="n">book</span> <span class="ow">in</span> <span class="nf">sorted</span><span class="p">(</span><span class="n">my_books</span><span class="p">):</span>
  <span class="n">output</span> <span class="o">=</span> <span class="sh">"</span><span class="s">{author}, {title}</span><span class="sh">"</span><span class="p">.</span><span class="nf">format</span><span class="p">(</span>
    <span class="n">author</span><span class="o">=</span><span class="n">book</span><span class="p">.</span><span class="n">author</span><span class="p">,</span>
    <span class="n">title</span><span class="o">=</span><span class="n">book</span><span class="p">.</span><span class="n">title</span><span class="p">,</span>
  <span class="p">)</span>
  <span class="nf">print</span><span class="p">(</span><span class="n">output</span><span class="p">)</span>

<span class="c1"># &gt;&gt; Ernest Hemingway, For Whom The Bell Tolls
# &gt;&gt; Ernest Hemingway, The Sun Also Rises
# &gt;&gt; William Faulkner, Absalom, Absalom!
# &gt;&gt; William Faulkner, The Sound and the Fury
</span></code></pre></div></div>

<p>And we didn’t have to write six methods!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:to">
      <p>A <a href="https://docs.python.org/3/glossary.html#term-decorator">decorator</a> is a function that takes a Python object as an argument and returns a (often) modified copy of the object. <a href="#fnref:to" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Your classes can make use of the rich Python comparison operators just like the built-in classes. Here I'll show you how to do it while minimizing boilerplate.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/patterns/biologia_centrali_americana_coronella_annulata.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/patterns/biologia_centrali_americana_coronella_annulata.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Data Science Salaries</title>
      <link href="https://alexgude.com/blog/data-science-salaries/" rel="alternate" type="text/html" title="Data Science Salaries" />
      <published>2019-03-26T00:00:00-07:00</published>
      <updated>2019-03-26T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/data_science_salaries</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/data-science-salaries/"><![CDATA[<p>One of the most important things to know when looking for a job is your market
value. For data science positions, what better way to determine that than with
data!<sup style="anchor-name:--fnref-data" id="fnref:data"><a href="#fn:data" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Some sites⁠—⁠like <a href="https://www.indeed.com/salaries/Data-Scientist-Salaries,-Mountain-View-CA">Indeed</a> and
<a href="https://www.glassdoor.com/Salaries/san-jose-data-scientist-salary-SRCH_IL.0,8_IM761_KO9,23.htm">Glassdoor</a>⁠—⁠offer aggregate salary information, but
they won’t give you all their data.<sup style="anchor-name:--fnref-levels" id="fnref:levels"><a href="#fn:levels" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> That’s why I prefer the survey that
<a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight</a> alumni put together. With the full dataset I can slice it
however I want.</p>

<p>The data is available <a href="/files/data-science-salaries//insight_salary_survey_cleaned.csv">here</a>. The notebook with all the code is
<a href="/files/data-science-salaries//Data%20Science%20Salary%20Data.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/data-science-salaries//Data%20Science%20Salary%20Data.ipynb">rendered on Github</a>).</p>

<h2 id="scientists-engineers-and-analysts">Scientists, Engineers, and Analysts</h2>

<p>The dataset is mostly comprised of data scientists, but there are a few
engineers and analysts as well, so we can compare salary across job titles:</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_vs_others.svg"><img src="/files/data-science-salaries//data_science_total_comp_vs_others.svg" alt="Box plot showing that machine learning engineers earn the most, followed by
data scientists." /></a></p>

<p>Machine learning engineers have the highest median salary, followed by data
scientists, data engineers, and data analysts. Data science salaries span a
greater range, although this is likely due to the larger number in the sample
(N=132 vs N=8). The median values are:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Title</th>
      <th style="text-align: right">Median Total Compensation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Machine Learning Engineer</td>
      <td style="text-align: right">$161k</td>
    </tr>
    <tr>
      <td style="text-align: left">Data Scientist</td>
      <td style="text-align: right">$145k</td>
    </tr>
    <tr>
      <td style="text-align: left">Data Engineer</td>
      <td style="text-align: right">$128k</td>
    </tr>
    <tr>
      <td style="text-align: left">Analyst</td>
      <td style="text-align: right">$105k</td>
    </tr>
  </tbody>
</table>

<h2 id="location-location-location">Location, Location, Location</h2>

<p>The market rate for a data scientist varies wildly by location:</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_by_region.svg"><img src="/files/data-science-salaries//data_science_total_comp_by_region.svg" alt="Box plot showing that the West Coast has the highest total compensation,
followed by New York City." /></a></p>

<p>Compensation in California is the highest, by far. Technology companies often
give employees stock in addition to their base pay and many of these companies
are based in California. Do those stock grants account for most of the
compensation difference between California and the rest of the nation? Let’s
remove stock and look again:</p>

<p><a href="/files/data-science-salaries//data_science_salary_by_region.svg"><img src="/files/data-science-salaries//data_science_salary_by_region.svg" alt="Box plot showing that the West Coast has the highest total compensation,
followed by New York City." /></a></p>

<p>The median compensation in California drops $20k when removing stock, the
median in the Northwest (another major tech hub) drops $5k, and the Northeast
and Midwest don’t change. The highest paid individuals in California all drop
drastically with the new highest paid person in New York!</p>

<h2 id="experience-counts-a-lot">Experience Counts (A Lot!)</h2>

<p>Finally, how much does experience matter? Looking at just California this
time, because I have the most data for salaries there:</p>

<p><a href="/files/data-science-salaries//data_science_total_comp_by_experience.svg"><img src="/files/data-science-salaries//data_science_total_comp_by_experience.svg" alt="Box plot showing each year of experience increases a data scientist's total
compensation tremendously." /></a></p>

<p>A lot! Going from one to three years of experience increases the median
compensation by $100k!</p>

<p>But what is happening to people with four years of experience? There are very
few in the survey (N=4), so the median can be pulled by any particular
outlier. For example, if one of them got a job in 2012 and stayed there for
four years, they would miss out on the large raises that come from jumping to
a new company. Additionally, one of the data scientists who said they had four
years experience also said they were in a “Junior” position, so it’s possible
they counted their school work when answering whereas most would not count
school.</p>

<p>I have found it is much easier to get a large raise when starting a new
position, so this plot argues for the importance of not staying at your first
job too long!</p>

<h2 id="now-you-know-but-knowing-is-only-half-the-battle">Now You Know (But Knowing Is Only Half the Battle)</h2>

<p>Now you have a better idea what the data science job market looks like, but
that isn’t enough. To get what you’re worth, you have to negotiate as well. I
highly recommend reading <a href="https://twitter.com/patio11">Patrick McKenzie’s</a> <a href="https://www.kalzumeus.com/2012/01/23/salary-negotiation/"><em><strong>Salary Negotiation:
Make More Money, Be More Valued</strong></em></a> post. Negotiating has earned me
between 5% and 10% increases to my offers, which as you can see from the
numbers in this post, are substantial! I wrote about my experience negotiating
with some tips here: <a href="/blog/data-science-asking-for-more-money/"><em><strong>Data Science, Compensation, and Asking for
Money</strong></em></a>.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:data">
      <p>Companies pay good money to get salary data so they know exactly what a good (and bad) offer looks like. You should have the same information. <a href="#fnref:data" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:levels">

      <p>One site, <a href="https://levels.fyi">levels.fyi</a>, <strong>does</strong> offer their data and has
very accurate total compensation information. Definitely give them a look if
you are trying to figure out what your skills are worth! <a href="#fnref:levels" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[How do data scientist salaries vary by experience and location? Read on to find out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-salaries/josef_wagner_hohenberg_the_billing_coins.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-salaries/josef_wagner_hohenberg_the_billing_coins.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: On What Days Do Cyclists Crash?</title>
      <link href="https://alexgude.com/blog/switrs-bicycle-crashes-by-date/" rel="alternate" type="text/html" title="SWITRS: On What Days Do Cyclists Crash?" />
      <published>2019-02-20T00:00:00-08:00</published>
      <updated>2019-02-20T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_bicycle_crashes_by_date</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-bicycle-crashes-by-date/"><![CDATA[<p>It is time to use <a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">SWITRS data</a> to look at vehicle crashes in
California again. I have previously used the data to look at <a href="/blog/switrs-crashes-by-date/">when cars
crash</a>—during holidays when people both drive to work and to
parties after—and <a href="/blog/switrs-motorcycle-crashes-by-date/">when motorcycles crash</a>—during the summer
when it’s good riding weather. Today I want to look at something a little
closer to my heart: <strong>bicycles</strong>.</p>

<p>I have been commuting on my bike for years now, and when I was younger I used
to put in thousands of miles a year for fun. So knowing more about when
crashes happen is something I am very interested in.</p>

<p>As per usual, the Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-bicycle-accidents-by-date/SWITRS%20Crash%20Dates%20With%20Bicycles.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-bicycle-accidents-by-date/SWITRS%20Crash%20Dates%20With%20Bicycles.ipynb">rendered on Github</a>).</p>

<h2 id="a-simple-model">A Simple Model</h2>

<p>Before we dig into the data, I have a simple model for how many bicycle
crashes there are. It is:</p>

<!-- prettier-ignore-start -->

\[N_{\textrm{crashes}} = P_{\textrm{car-bike}} \, L_{\textrm{miles biked}} \, \lambda_{\textrm{cars per mile}}\]

<!-- prettier-ignore-end -->

<p>That is, the number of crashes involving bicycles (\(N\)) is the probability
of a crash happening when a bike encounters a car (\(P\)) times the number of
cars encountered (\(L \lambda\)). This ignores some crashes, like solo crashes
and those that do not involve a car, but these are rare.<sup style="anchor-name:--fnref-rare" id="fnref:rare"><a href="#fn:rare" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>We won’t be able to test the validity of this model with the SWITRS data
alone, but we can use it to reason about what is happening. For example, if
the number of crashes increases, that could be because there are more cars or
bikes on the road, or because the probability of collision increased (perhaps
due to distracted drivers or worse average weather).</p>

<h2 id="data-selection">Data Selection</h2>

<p>I selected crashes involving bicycles from the <a href="https://github.com/agude/SWITRS-to-SQLite">SQLite database</a>
(<a href="/blog/switrs-to-sqlite/">discussed previously</a>) with the following query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">Collision_Date</span> <span class="k">FROM</span> <span class="n">Collision</span>
<span class="k">WHERE</span> <span class="n">Collision_Date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">Bicycle_Collision</span> <span class="o">==</span> <span class="mi">1</span>          <span class="c1">-- Involves a bicycle</span>
<span class="k">AND</span> <span class="n">Collision_Date</span> <span class="o">&lt;=</span> <span class="s1">'2017-12-31'</span>  <span class="c1">-- 2018 is incomplete</span>
</code></pre></div></div>

<p>This gave me 223,772 data points to examine spanning 2001 to 2017. <a href="/blog/switrs-crashes-by-date/#data-selection">Just as
before</a>, crashes from the most recent year are rejected because the
database dump comes from September 2018, and so the year is incomplete.</p>

<h2 id="crashes-per-week">Crashes per Week</h2>

<p>For car crashes, <a href="/blog/switrs-crashes-by-date/#crashes-per-week">I found that there was a large dip in 2008</a> as
people stopped driving to work during the <a href="https://en.wikipedia.org/wiki/Great_Recession">Great Recession</a>. For
motorcycle crashes, <a href="/blog/switrs-motorcycle-crashes-by-date/#crashes-per-week">I found strong seasonality</a> as people hung up
their helmets during the winter. For bicycles, we have the following pattern:</p>

<p><a href="/files/switrs-bicycle-accidents-by-date/bicycle_accidents_per_week_in_california.svg"><img src="/files/switrs-bicycle-accidents-by-date/bicycle_accidents_per_week_in_california.svg" alt="Line plot showing bicycles crashes per week from 2001 through
2017" /></a></p>

<p>It shows features similar to both cars and motorcycles:</p>

<ul>
  <li>
    <p>The number of crashes increases after 2008, and then begins decreasing after
2013, almost exactly the <strong>opposite</strong> of the car pattern.</p>
  </li>
  <li>
    <p>Crashes are highly seasonal, just like motorcycles. Apparently neither
cyclists nor bikers like riding in the rain.</p>
  </li>
</ul>

<p>Thinking back to <a href="#a-simple-model">the model</a> we can try to reason about the trend. We
know the number of cars increased, so the decrease in crashes in the last few
years is either due to a decrease in the number of cyclists—possibly because
they traded their bikes for cars as they found employment—or a decrease in
the likelihood of a crashes—perhaps because drivers are more used to
cyclists and look out for them.</p>

<h2 id="day-by-day">Day-by-Day</h2>

<p>Cars are involved in crashes <a href="/blog/switrs-crashes-by-date/#day-by-day">on holidays during which the drivers also
work</a>, like Halloween. Motorcycles are in crashes <a href="/blog/switrs-motorcycle-crashes-by-date/#day-by-day">during summer
holidays</a>. Bicycles, on the other hand, have no holidays with a large
excess in the number of crashes. Some holidays, like Christmas and
Thanksgiving, keep people from getting on their bikes, but none seem to
motivate people to get out and ride.</p>

<p><a href="/files/switrs-bicycle-accidents-by-date/mean_bicycle_accidents_by_date.svg"><img src="/files/switrs-bicycle-accidents-by-date/mean_bicycle_accidents_by_date.svg" alt="Line plot showing average motorcycle crashes by day of the
year" /></a></p>

<p>New Year’s Day, St. Patrick’s Day, and the 4th of July are all higher than
they would be if they were not holidays, although you can’t tell from this
plot. On those days, people tend to go out and celebrate with alcohol, which
leads to solo crashes. I will examine that in a future post.</p>

<h2 id="day-of-the-week">Day of the Week</h2>

<p>For cars, <a href="/blog/switrs-crashes-by-date/#day-of-the-week">weekends show a decrease in the number of crashes</a> as
people stop commuting. For motorcycles, <a href="/blog/switrs-motorcycle-crashes-by-date/#day-of-the-week">weekends show an increase in the
number of crashes</a> as people use their time off to ride. As a
recreational cyclist, I expected crashes to increase on the weekend as people
put on their Lycra and take to the back roads for fun. But this is not the
case:</p>

<p><a href="/files/switrs-bicycle-accidents-by-date/bicycle_accidents_by_day_of_the_week.svg"><img src="/files/switrs-bicycle-accidents-by-date/bicycle_accidents_by_day_of_the_week.svg" alt="Violin plot showing the number of bicycle crashes by day of the
week" /></a></p>

<p>These <a href="https://en.wikipedia.org/wiki/Violin_plot">violin plots</a> show the distribution of crashes by day of the
week over the 17 year period. There is a large drop in the number of crashes
on weekends. This is surprising to me. I would have expected a lot more
cyclists to be out on the weekend, leading to more interactions with cars.</p>

<p>It’s possible that there are more cyclists on the weekend but there are enough
fewer cars that the crash rate still goes down. Or perhaps the riders are
better at avoiding crashes. Or maybe the cyclists are out in the countryside
away from the cars. Or perhaps weekend drivers are better at avoiding
cyclists. Without more data, we can’t tell.</p>

<h2 id="conclusion">Conclusion</h2>

<p>This analysis of bicycle crashes surprised me a little. I expected bikes to
show a similar pattern to motorcycles, since they are both used to commute and
for fun. However, bikes show a greatly reduced crash rate on the weekend while
motorcycles show an increase. Bikes and cars also seem to trade off, with car
crashes increasing in recent years while bike crashes fall off. Further study
and additional data is necessary before I can determine the reasons behind
this trend.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:rare">

      <p>Of the 223,772 recorded crashes with bicycles, <strong>89% involve a car</strong>.
There is a bias though: SWITRS reports are filled out when the police or
CHP are called to the scene. As such, they skew towards worse accidents. <a href="#fnref:rare" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="cycling" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[California crash data doesn't just cover cars, it covers bikes too! This time we look at when cyclists crash in California.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-bicycle-accidents-by-date/wilhelmina_cycle_co.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-bicycle-accidents-by-date/wilhelmina_cycle_co.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Python Patterns: Enum</title>
      <link href="https://alexgude.com/blog/python-patterns-enum/" rel="alternate" type="text/html" title="Python Patterns: Enum" />
      <published>2019-01-22T00:00:00-08:00</published>
      <updated>2019-01-22T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/python_patterns_enum</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-patterns-enum/"><![CDATA[<p>Things often come in sets, for example, States, Pokémon, Playing cards, etc.</p>

<p>Each set has items that belong to them (like California, Charizard, Jack of
Clubs) and checking if an item is a valid member is a common task. Some
collections (playing cards, for example) are also orderable; twos come before
fives which come before Kings.</p>

<p>There are many ways to represent members from these sets in Python:</p>

<ul>
  <li>
    <p>Unique string: <code class="language-plaintext highlighter-rouge">"CA"</code>, <code class="language-plaintext highlighter-rouge">"WA"</code>, <code class="language-plaintext highlighter-rouge">"MN"</code></p>
  </li>
  <li>
    <p>Classes: <code class="language-plaintext highlighter-rouge">class Pokemon: ... </code></p>
  </li>
  <li>
    <p>Tuples (or <a href="/blog/python-patterns-namedtuple/">namedtuples</a>): <code class="language-plaintext highlighter-rouge">("Clubs", "J")</code>, <code class="language-plaintext highlighter-rouge">("Hearts", 5)</code></p>
  </li>
</ul>

<p>But is <code class="language-plaintext highlighter-rouge">"PR"</code> a valid state? Is <code class="language-plaintext highlighter-rouge">Pokemon("Digimon")</code> a member of Pokemon? Is
<code class="language-plaintext highlighter-rouge">("Lotus", "Black")</code> a playing card? We could keep a separate <code class="language-plaintext highlighter-rouge">set()</code> of all
valid members to check, but then we have to maintain it.</p>

<p>Thankfully, Python provides a way to create these sets and their members at
the same time: <a href="https://docs.python.org/3/library/enum.html"><strong>enumerations</strong></a>, or enums.</p>

<h2 id="playing-cards">Playing Cards</h2>

<p>Without using enums we might implement a <a href="https://en.wikipedia.org/wiki/Standard_52-card_deck">standard playing card</a> as
follows:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@total_ordering</span>
<span class="k">class</span> <span class="nc">PlayingCard</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">suit</span><span class="p">,</span> <span class="n">rank</span><span class="p">):</span>
    <span class="n">self</span><span class="p">.</span><span class="n">suit</span> <span class="o">=</span> <span class="n">suit</span>
    <span class="n">self</span><span class="p">.</span><span class="n">rank</span> <span class="o">=</span> <span class="n">rank</span>
    <span class="n">self</span><span class="p">.</span><span class="nf">__rank_to_value</span><span class="p">()</span>

  <span class="k">def</span> <span class="nf">__rank_to_value</span><span class="p">(</span><span class="n">self</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s"> Convert face cards to integer values. </span><span class="sh">"""</span>
    <span class="k">if</span> <span class="n">rank</span> <span class="o">==</span> <span class="sh">"</span><span class="s">A</span><span class="sh">"</span><span class="p">:</span>
      <span class="n">self</span><span class="p">.</span><span class="n">__value</span> <span class="o">=</span> <span class="mi">14</span>
    <span class="k">elif</span> <span class="n">rank</span> <span class="o">==</span> <span class="sh">"</span><span class="s">K</span><span class="sh">"</span><span class="p">:</span>
      <span class="n">self</span><span class="p">.</span><span class="n">__value</span> <span class="o">=</span> <span class="mi">13</span>
    <span class="p">...</span> <span class="c1"># etc.
</span>    <span class="c1"># Numbered cards are easy
</span>    <span class="k">else</span><span class="p">:</span>
      <span class="n">self</span><span class="p">.</span><span class="n">__value</span> <span class="o">=</span> <span class="n">self</span><span class="p">.</span><span class="n">rank</span>

  <span class="k">def</span> <span class="nf">__lt__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">__value</span> <span class="o">&lt;</span> <span class="n">other</span><span class="p">.</span><span class="n">__value</span>

  <span class="k">def</span> <span class="nf">__eq__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">__value</span> <span class="o">==</span> <span class="n">other</span><span class="p">.</span><span class="n">__value</span>
</code></pre></div></div>

<p>This class works with the standard comparison operators (thanks to <a href="/blog/python-patterns-total-ordering/">the
<code class="language-plaintext highlighter-rouge">@total_ordering</code> decorator, which I discuss in another
post</a>), but to do so we had to write a bit of an annoying
<code class="language-plaintext highlighter-rouge">__rank_to_value()</code> function; otherwise Aces and Kings would be tough to
compare to 2s and 3s!</p>

<p>With that done, we can now declare cards easily enough:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">ace_of_spades</span>    <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Spade</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">A</span><span class="sh">"</span><span class="p">)</span>
<span class="n">king_of_hearts</span>   <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Heart</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">K</span><span class="sh">"</span><span class="p">)</span>
<span class="n">eight_of_spades</span>  <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Spades</span><span class="sh">"</span><span class="p">,</span> <span class="mi">8</span><span class="p">)</span>
<span class="n">eight_of_clubs</span>   <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Club</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">8</span><span class="sh">"</span><span class="p">)</span>
<span class="n">my_favorite_card</span> <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Stars</span><span class="sh">"</span><span class="p">,</span> <span class="mi">85</span><span class="p">)</span>
</code></pre></div></div>

<p>Did you catch all the errors? We could write some error checking in the class,
but it would again be a bit tedious. Instead, let’s implement this using enums.
Enums will let us represent the suits and ranks, check that they are valid,
and order based on value, without writing a lot of extra code.</p>

<h2 id="playing-cards-with-enums">Playing Cards with Enums</h2>

<p>An enum has exactly the properties we want:</p>

<ul>
  <li>
    <p>We can test membership, so only real suits and ranks are allowed.</p>
  </li>
  <li>
    <p>We can order them, so we know that King &gt; Jack &gt; Ten.</p>
  </li>
</ul>

<p>First, we define the suits:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">enum</span> <span class="kn">import</span> <span class="n">Enum</span><span class="p">,</span> <span class="n">auto</span>

<span class="k">class</span> <span class="nc">CardSuit</span><span class="p">(</span><span class="n">Enum</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s"> Playing card suits. </span><span class="sh">"""</span>
    <span class="n">CLUBS</span> <span class="o">=</span> <span class="nf">auto</span><span class="p">()</span>
    <span class="n">DIAMONDS</span> <span class="o">=</span> <span class="nf">auto</span><span class="p">()</span>
    <span class="n">HEARTS</span> <span class="o">=</span> <span class="nf">auto</span><span class="p">()</span>
    <span class="n">SPADES</span> <span class="o">=</span> <span class="nf">auto</span><span class="p">()</span>
</code></pre></div></div>

<p>The function <code class="language-plaintext highlighter-rouge">auto()</code> sets the values and ensures that they are unique. The
members are not orderable (so <code class="language-plaintext highlighter-rouge">CardSuit.CLUBS &gt; CardSuit.DIAMONDS</code> will raise
an error), but do have equality (so <code class="language-plaintext highlighter-rouge">CardSuit.CLUBS != CardSuit.DIAMONDS</code>
works). We can also test membership easily, allowing us to ensure only valid
suits are accepted.</p>

<p>An example of some of the properties:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">hearts</span> <span class="o">=</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span>
<span class="n">clubs</span> <span class="o">=</span> <span class="n">CardSuit</span><span class="p">.</span><span class="n">CLUBS</span>
<span class="n">stars</span> <span class="o">=</span> <span class="sh">"</span><span class="s">stars</span><span class="sh">"</span>

<span class="c1"># We can test equality
</span><span class="n">hearts</span> <span class="o">!=</span> <span class="n">stars</span>  <span class="c1"># True
</span>
<span class="c1"># And we can test membership
</span><span class="nf">isinstance</span><span class="p">(</span><span class="n">stars</span><span class="p">,</span> <span class="n">CardSuit</span><span class="p">)</span>  <span class="c1"># False
</span></code></pre></div></div>

<p>Second, we define a <code class="language-plaintext highlighter-rouge">CardValue</code>, this time using <code class="language-plaintext highlighter-rouge">IntEnum</code> because we want the
values to be comparable.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">enum</span> <span class="kn">import</span> <span class="n">IntEnum</span><span class="p">,</span> <span class="n">unique</span>

<span class="nd">@unique</span>
<span class="k">class</span> <span class="nc">CardRank</span><span class="p">(</span><span class="n">IntEnum</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s"> Playing card values. They are orderable as expected:
    2 &lt; 3 &lt; ... &lt; king &lt; ace.
    </span><span class="sh">"""</span>
    <span class="n">TWO</span> <span class="o">=</span> <span class="mi">2</span>
    <span class="n">THREE</span> <span class="o">=</span> <span class="mi">3</span>
    <span class="p">...</span> <span class="c1"># etc.
</span>    <span class="n">TEN</span> <span class="o">=</span> <span class="mi">10</span>
    <span class="n">JACK</span> <span class="o">=</span> <span class="mi">11</span>
    <span class="n">QUEEN</span> <span class="o">=</span> <span class="mi">12</span>
    <span class="n">KING</span> <span class="o">=</span> <span class="mi">13</span>
    <span class="n">ACE</span> <span class="o">=</span> <span class="mi">14</span>
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">IntEnum</code>s are orderable so <code class="language-plaintext highlighter-rouge">CardRank.TEN &lt; CardRank.KING</code>. The decorator
<code class="language-plaintext highlighter-rouge">@unique</code> adds a check that makes sure we haven’t double assigned any values.</p>

<p>Now the card class is easy to implement:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@total_ordering</span>
<span class="k">class</span> <span class="nc">PlayingCard</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">suit</span><span class="p">,</span> <span class="n">rank</span><span class="p">):</span>
    <span class="c1"># Check that the suit is valid
</span>    <span class="k">if</span> <span class="ow">not</span> <span class="nf">isinstance</span><span class="p">(</span><span class="n">suit</span><span class="p">,</span> <span class="n">CardSuit</span><span class="p">):</span>
      <span class="k">raise</span> <span class="nc">ValueError</span><span class="p">(</span><span class="sh">"</span><span class="s">{} is an invalid CardSuit.</span><span class="sh">"</span><span class="p">.</span><span class="nf">format</span><span class="p">(</span><span class="n">suit</span><span class="p">))</span>
    <span class="n">self</span><span class="p">.</span><span class="n">suit</span> <span class="o">=</span> <span class="n">suit</span>

    <span class="c1"># Check that the rank is valid
</span>    <span class="k">if</span> <span class="ow">not</span> <span class="nf">isinstance</span><span class="p">(</span><span class="n">rank</span><span class="p">,</span> <span class="n">CardRank</span><span class="p">):</span>
      <span class="k">raise</span> <span class="nc">ValueError</span><span class="p">(</span><span class="sh">"</span><span class="s">{} is an invalid CardRank.</span><span class="sh">"</span><span class="p">.</span><span class="nf">format</span><span class="p">(</span><span class="n">rank</span><span class="p">))</span>
    <span class="n">self</span><span class="p">.</span><span class="n">rank</span> <span class="o">=</span> <span class="n">rank</span>

  <span class="k">def</span> <span class="nf">__lt__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">rank</span> <span class="o">&lt;</span> <span class="n">other</span><span class="p">.</span><span class="n">rank</span>

  <span class="k">def</span> <span class="nf">__eq__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">other</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">rank</span> <span class="o">==</span> <span class="n">other</span><span class="p">.</span><span class="n">rank</span>
</code></pre></div></div>

<p>It is now much easier to catch errors in our card definitions:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">ace_of_spades</span>    <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="n">CardSuit</span><span class="p">.</span><span class="n">SPADES</span><span class="p">,</span> <span class="n">CardRank</span><span class="p">.</span><span class="n">ACE</span><span class="p">)</span>
<span class="n">king_of_hearts</span>   <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="n">CardSuit</span><span class="p">.</span><span class="n">HEARTS</span><span class="p">,</span> <span class="n">CardRank</span><span class="p">.</span><span class="n">KING</span><span class="p">)</span>
<span class="n">my_favorite_card</span> <span class="o">=</span> <span class="nc">PlayingCard</span><span class="p">(</span><span class="sh">"</span><span class="s">Stars</span><span class="sh">"</span><span class="p">,</span> <span class="mi">85</span><span class="p">)</span>  <span class="c1"># Obviously wrong
</span></code></pre></div></div>

<p>Not only are they obvious by eye (<code class="language-plaintext highlighter-rouge">"Stars"</code> is clearly not a <code class="language-plaintext highlighter-rouge">CardSuit</code>), but
the runtime will even raise an error!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Things often come in sets of specific items, like states, Pokémon, or playing cards. Python has an elegant way of representing them using enum.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/patterns/locupletissimi_rerum_naturalium_thesauri_v1_lxxxiii_snake.png" />
        <media:content medium="image" url="https://alexgude.com/files/patterns/locupletissimi_rerum_naturalium_thesauri_v1_lxxxiii_snake.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Python Patterns: Named Tuples</title>
      <link href="https://alexgude.com/blog/python-patterns-namedtuple/" rel="alternate" type="text/html" title="Python Patterns: Named Tuples" />
      <published>2018-12-18T00:00:00-08:00</published>
      <updated>2018-12-18T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/python_patterns_namedtuple</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-patterns-namedtuple/"><![CDATA[<p>In Python, <a href="https://docs.python.org/3.7/library/stdtypes.html#sequence-types-list-tuple-range">sequences</a> are a great way to hold a set of ordered data. As
long as the data is simple, a list or tuple is perfect because they are
included in every install of Python. But data is not always simple; you can
put any object you want in a sequence, making it easy to lose track of what is
where.</p>

<p>For example, one might create cards in a virtual address book like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">card</span> <span class="o">=</span> <span class="p">(</span>
  <span class="sh">"</span><span class="s">Alex</span><span class="sh">"</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">Gude</span><span class="sh">"</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">me@alexgude.com</span><span class="sh">"</span><span class="p">,</span>
  <span class="bp">None</span><span class="p">,</span>
  <span class="sh">"</span><span class="s">17 St., Smaller Town, CA</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>
</code></pre></div></div>

<p>Simple, but a little confusing. What does <code class="language-plaintext highlighter-rouge">None</code> signify? Writing code to work
with these objects is error prone:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">check_email</span><span class="p">(</span><span class="n">card</span><span class="p">):</span>
  <span class="sh">"""</span><span class="s">Check if a card has an email
  address that is valid.</span><span class="sh">"""</span>
  <span class="n">email</span> <span class="o">=</span> <span class="n">card</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span>  <span class="c1"># 2?!
</span>  <span class="n">is_valid</span> <span class="o">=</span> <span class="n">email</span> <span class="ow">is</span> <span class="ow">not</span> <span class="bp">None</span> <span class="ow">and</span> <span class="sh">'</span><span class="s">@</span><span class="sh">'</span> <span class="ow">in</span> <span class="n">email</span>

  <span class="k">return</span> <span class="n">is_valid</span>
</code></pre></div></div>

<p>Is <code class="language-plaintext highlighter-rouge">2</code> the correct index to use? Maybe it was <code class="language-plaintext highlighter-rouge">3</code>? Catching mistakes in the
code is tough for anyone reading it.</p>

<h2 id="alternatives">Alternatives</h2>

<p>A dictionary is a natural solution to this problem, because we can use strings
as keys, for example <code class="language-plaintext highlighter-rouge">card["email"]</code> instead of <code class="language-plaintext highlighter-rouge">card[2]</code>. But we might need
to maintain compatibility with something that expects a sequence, as was the
case when <a href="/blog/matplotlib-blitting-supernova/#blitting">passing artists around in my <code class="language-plaintext highlighter-rouge">matplotlib</code> blitting post</a>.</p>

<p>Instead, we could build a class that acts like a list or tuple::</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Card</span><span class="p">:</span>
  <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">first_name</span><span class="p">,</span> <span class="n">last_name</span><span class="p">,</span> <span class="p">...):</span>
    <span class="n">self</span><span class="p">.</span><span class="n">__internal</span> <span class="o">=</span> <span class="p">[</span><span class="n">first_name</span><span class="p">,</span> <span class="n">last_name</span><span class="p">,</span> <span class="p">...]</span>
    <span class="n">self</span><span class="p">.</span><span class="n">first_name</span> <span class="o">=</span> <span class="n">self</span><span class="p">.</span><span class="n">__internal</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
    <span class="n">self</span><span class="p">.</span><span class="n">last_name</span> <span class="o">=</span> <span class="n">self</span><span class="p">.</span><span class="n">__internal</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
    <span class="p">...</span>  <span class="c1"># etc.
</span>
  <span class="k">def</span> <span class="nf">__len__</span><span class="p">(</span><span class="n">self</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">__internal</span><span class="p">.</span><span class="nf">__len__</span><span class="p">()</span>

  <span class="k">def</span> <span class="nf">__getitem__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">key</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">__internal</span><span class="p">.</span><span class="nf">__getitem__</span><span class="p">(</span><span class="n">key</span><span class="p">)</span>

  <span class="k">def</span> <span class="nf">__next__</span><span class="p">(</span><span class="n">self</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">self</span><span class="p">.</span><span class="n">__internal</span><span class="p">.</span><span class="nf">__next__</span><span class="p">()</span>

  <span class="c1"># and many other methods
</span></code></pre></div></div>

<p>Not difficult to write, but tedious due to all of the boilerplate code.
Thankfully, someone has already done so.</p>

<h2 id="named-tuples">Named Tuples</h2>

<p>The <a href="https://docs.python.org/3/library/collections.html#collections.namedtuple">named tuple</a> functions exactly like a tuple, with one
addition: you can access each component of the tuple by name. Our card example
would now look like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">collections</span> <span class="kn">import</span> <span class="n">namedtuple</span>

<span class="n">Card</span> <span class="o">=</span> <span class="nf">namedtuple</span><span class="p">(</span>
    <span class="sh">"</span><span class="s">Card</span><span class="sh">"</span><span class="p">,</span>
    <span class="p">[</span>
        <span class="sh">"</span><span class="s">first_name</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">last_name</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">email</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">phone</span><span class="sh">"</span><span class="p">,</span>  <span class="c1"># Our empty field revealed!
</span>        <span class="sh">"</span><span class="s">address</span><span class="sh">"</span><span class="p">,</span>
    <span class="p">]</span>

<span class="p">)</span>

<span class="n">alex_card</span> <span class="o">=</span> <span class="nc">Card</span><span class="p">(</span>
    <span class="sh">"</span><span class="s">Alex</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">Gude</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">me@alexgude.com</span><span class="sh">"</span><span class="p">,</span>
    <span class="bp">None</span><span class="p">,</span> <span class="sh">"</span><span class="s">17 St., Smaller Town, CA</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>
</code></pre></div></div>

<p>This is much cleaner than our original card tuple. We now know the missing
value is the phone number! We can access the values with dot operators as
well: <code class="language-plaintext highlighter-rouge">card.email</code>. And the named tuple stills works exactly as you would
expect for a standard tuple:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># For loops work
</span><span class="k">for</span> <span class="n">item</span> <span class="ow">in</span> <span class="n">alex_card</span><span class="p">:</span>
    <span class="nf">print</span><span class="p">(</span><span class="n">item</span><span class="p">)</span>

<span class="c1"># We can access with . or []
</span><span class="n">alex_card</span><span class="p">[</span><span class="mi">2</span><span class="p">]</span> <span class="o">==</span> <span class="n">alex_card</span><span class="p">.</span><span class="n">email</span>

<span class="c1"># And we can unpack
</span><span class="n">first</span><span class="p">,</span> <span class="n">last</span><span class="p">,</span> <span class="n">email</span><span class="p">,</span> <span class="n">phone</span><span class="p">,</span> <span class="n">address</span> <span class="o">=</span> <span class="n">alex_tuple</span>
</code></pre></div></div>

<p>Code that operates on this named tuple is much easier to read as well, because
it does not rely on <a href="https://en.wikipedia.org/wiki/Magic_number_(programming)">magic numbers</a>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">new_check_email</span><span class="p">(</span><span class="n">card</span><span class="p">):</span>
  <span class="sh">"""</span><span class="s">Check if a card has an email
  address that is valid.</span><span class="sh">"""</span>
  <span class="n">email</span> <span class="o">=</span> <span class="n">card</span><span class="p">.</span><span class="n">email</span>
  <span class="n">is_valid</span> <span class="o">=</span> <span class="n">email</span> <span class="ow">is</span> <span class="ow">not</span> <span class="bp">None</span> <span class="ow">and</span> <span class="sh">'</span><span class="s">@</span><span class="sh">'</span> <span class="ow">in</span> <span class="n">email</span>

  <span class="k">return</span> <span class="n">is_valid</span>
</code></pre></div></div>

<p>Named tuples are not as well known as dictionaries or classes, but they solve
a common problem and make your code more readable!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Sometimes I need to store an ordered dataset, but reference specific members from it. Named tuples in Python provide a clean way to do this!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/patterns/lycodon_modestus.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/patterns/lycodon_modestus.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: Car Crashes After Daylight Saving Time Ends</title>
      <link href="https://alexgude.com/blog/switrs-daylight-saving-time-end-accidents/" rel="alternate" type="text/html" title="SWITRS: Car Crashes After Daylight Saving Time Ends" />
      <published>2018-11-03T00:00:00-07:00</published>
      <updated>2018-11-03T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/switrs_daylight_saving_time_end_accidents</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-daylight-saving-time-end-accidents/"><![CDATA[<p>We all hate the change to <a href="https://en.wikipedia.org/wiki/Daylight_saving_time">daylight saving time</a> (DST) in the spring; it
makes us tired, grumpy, but worst of all it <a href="/blog/switrs-daylight-saving-time-accidents/">causes us to crash our cars at a
higher rate</a>! The end of DST is not as universally reviled,
probably because we get back the hour of sleep we lost earlier in the year,
but <a href="https://doi.org/10.1016/S1389-9457(00)00032-0">Varughese &amp; Allen</a> found that there was still a “significant
increase in number of crashes on the Sunday of the fall shift from
DST”.<sup style="anchor-name:--fnref-varughese_cite" id="fnref:varughese_cite"><a href="#fn:varughese_cite" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>With the <a href="/blog/switrs-to-sqlite/">SWITRS data</a> that I collected, and the analysis code I
developed for <a href="/blog/switrs-daylight-saving-time-accidents/">my post last year looking at car crashes after the DST
change</a>, it should be pretty easy to check if I see the same
trend as Varughese &amp; Allen.</p>

<p>The Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-dst/SWITRS%20Daylight%20Saving%20Time%20Crashes.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-dst/SWITRS%20Daylight%20Saving%20Time%20Crashes.ipynb">rendered on Github</a>).</p>

<h2 id="crash-ratio">Crash Ratio</h2>

<p>Just like <a href="/blog/switrs-daylight-saving-time-accidents/#crash-ratio">last time</a>, I will look at the number of crashes on the
days following the end of DST. In order to help cancel out effects other than
the time change—like the fact that <a href="/blog/switrs-crashes-by-date/#crashes-per-week">crash rates vary by 30% depending on the
year</a>—I will divide each day’s total by the number of crashes a week
later, when people are presumably back to normal.</p>

<p>Unlike last time, I am not normalizing by the number of crashes two weeks
after the change. The reason for this is simple: that’s Thanksgiving week, and
<a href="/blog/switrs-crashes-by-date/#day-by-day">as I showed before</a> the number of crashes is greatly reduced
during the holidays.</p>

<p>Just as last time, the <a href="https://en.wikipedia.org/wiki/Violin_plot">violin plots</a> below show the distribution of
these ratios from the years 2001 to 2017. A value greater than 1 means that
there are more crashes during the week when DST ends than the week after.</p>

<p><a href="/files/switrs-dst/accidents_after_end_dst_in_california.svg"><img src="/files/switrs-dst/accidents_after_end_dst_in_california.svg" alt="Violin plot showing the ratio of crashes per day of the week for the week
daylight saving time ends, divided by the week
after." /></a></p>

<p>There is, on average, a larger number of crashes on Sunday when the time
changes as seen by Varughese &amp; Allen. However, the same excess is not seen
when a different normalization is chosen, like using the <a href="/files/switrs-dst/accidents_after_end_dst_in_california_before.svg">week
before</a> or <a href="/files/switrs-dst/accidents_two_weeks_after_end_dst_in_california.svg">two weeks after</a>.<sup style="anchor-name:--fnref-after" id="fnref:after"><a href="#fn:after" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> The week before
has different lighting during commute times and so it is easier to dismiss,
but two weeks after has similar lighting.</p>

<h2 id="t-test"><em>t</em>-Test</h2>

<p>Instead, we turn away from our “<em>chi-by-eye</em>” test and do an actual
statistical test: a <a href="https://en.wikipedia.org/wiki/Student%27s_t-test#Paired_samples"><em>two-tailed paired t-test</em></a>, the same test
used by Varughese &amp; Allen. They find a significant (<em>p</em> &lt; 0.002) increase in
the number of deadly crashes on the Sunday that DST ends, but I do not (<em>p</em>
= 0.082).</p>

<p>Our methods are different in a few key ways:</p>

<ul>
  <li>
    <p>They look only at fatal crashes while I look at all.</p>
  </li>
  <li>
    <p>They compare to the mean of the week before and after while I use only the week after.</p>
  </li>
</ul>

<p>If I reproduce their methods with my dataset, I still do not find a
significant result (<em>p</em> = 0.158).</p>

<p>As for California, <a href="https://en.wikipedia.org/wiki/Kansen_Chu">Kansen Chu</a> has once again given us a chance to get
rid of the time change with <a href="https://ballotpedia.org/California_Proposition_7,_Permanent_Daylight_Saving_Time_Measure_(2018)">Prop 7</a>. His <a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201520160AB385">earlier bill
failed</a>, so he has gone directly to the voters this time. Although
<a href="https://medium.com/@USC/why-proposition-7-is-bad-for-public-health-825905ba54f6">permanent DST is not the ideal solution</a>, I’m still for getting rid of
the time change itself!</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:varughese_cite">

      <p><span class="citation">J. Varughese and R. Allen. <a href="https://doi.org/10.1016/S1389-9457(00)00032-0">“Fatal accidents following changes in daylight savings time: the American experience”</a> <cite>Sleep Medicine</cite>. vol. 2, no. 1. 2000. pp. 31–36. doi: <a href="https://doi.org/10.1016/S1389-9457(00)00032-0">10.1016/S1389-9457(00)00032-0</a>.</span> <a href="#fnref:varughese_cite" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:after">
      <p>The large deviations on the two week plot for Thursday and Friday are explained by Thanksgiving and Black Friday. <a href="#fnref:after" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="california-traffic-data" />
        
      

      

      
      
        <summary type="html"><![CDATA[Daylight saving time leads to more traffic collisions, but what about when DST ends? Some researchers have found that it does lead to more crashes, so I take a look using California's SWITRS data.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-dst/a_woman_sets_the_clocks_forward.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-dst/a_woman_sets_the_clocks_forward.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Academic Bully at CERN</title>
      <link href="https://alexgude.com/blog/my-academic-bully/" rel="alternate" type="text/html" title="My Academic Bully at CERN" />
      <published>2018-10-15T00:00:00-07:00</published>
      <updated>2018-10-15T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/my_academic_bully</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-academic-bully/"><![CDATA[<p>Working at CERN was the highlight of my academic career. In many ways it is
the perfect place for a scientist, a place where you get to work on the most
cutting-edge problems in your field while surrounded by people who all share
the same passion for the work. But in one way, for me, it was horrible: it is
where I met the woman who would abuse me until I abandoned the project I had been
sent over to work on.</p>

<p>I am not the kind of person who I thought would be the target of <a href="https://doi.org/10.1038/d41586-018-06040-w">academic
bullying</a>. I’m big, I’m tough, I get along with pretty much
everyone. My abuser did not match my stereotypes of a bully either: she was a
woman, diminutive, seemingly unthreatening.</p>

<p>Writing this post, even years after the fact, was surprisingly tough; just
thinking back on the events and the person raises my heart rate and makes me
feel anxious and uneasy. But I felt like I had to write it, for me, and for
everyone else who has gone through something similar. I hope that if enough
people speak up, we will eventually be able to fight the abuse instead of
hiding it.</p>

<h2 id="heading-to-cern-and-the-start">Heading to CERN, and The Start</h2>

<p>In 2012, I went to <a href="https://home.cern/">CERN</a> to complete some of my required service work
on the <a href="https://en.wikipedia.org/wiki/Compact_Muon_Solenoid">Compact Muon Solenoid experiment</a>. Service work is the mopping
and sweeping of experimental physics: it must be done to produce the desired
scientific results, but it is not really science in and of itself. For my
service work, I was a High Level Trigger (HLT) on-call. I took rotations
monitoring the system and I had to be available to troubleshoot problems at a
moment’s notice, day or night.</p>

<p>The HLT group ran the on-call program. The group was composed of grad students
and post docs (and some professors) who wrote the software, built the
hardware, and monitored the system. The post doc I worked with from Minnesota
ran part of the group, and his deputy was a French post doc and my eventual
abuser. She did not like me from the get-go; I suspect she saw me as a natural
ally to my post doc, with whom she was in competition.<sup style="anchor-name:--fnref-tenure" id="fnref:tenure"><a href="#fn:tenure" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>The HLT post docs and grad students spent lots of time training me to handle
the responsibilities of the on-call position. In fact, as one of the first
people to go through the training,<sup style="anchor-name:--fnref-training" id="fnref:training"><a href="#fn:training" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> they spent extra time with me as
any misunderstanding on my part indicated a problem with the onboarding
system. Although she would later abuse me for incompetence, my abuser refused
to contribute to my training.</p>

<h2 id="the-attacks-and-my-panic-attack">The Attacks and My Panic Attack</h2>

<p>I do not remember how the abuse started, but it soon became a constant
barrage. Emails sent to the entire HLT list telling me that I was an idiot.
Comments in group meetings saying the same. Saying that I was doing a terrible
job. That I was doing the work too slow. That I knew nothing. That I made them
look bad.</p>

<p>My panic attack came on suddenly. If asked, I would have guessed that a panic
attack must start with a moment of sheer panic, but mine did not. Mine started
with anger. I had just gotten back to my office from a meeting and opened my
email to find another one from my bully. I clenched my fists, imagining the
cutting response I would never write to shame her and make it stop. And then
it happened: I could not stop clenching my fists, or even move my arms. I sunk
to the floor as my heart raced in fear at having lost control of my movements.</p>

<p>My post doc came back to the office then. He helped me into our Peugeot
Partner CERN van and drove me to the hospital. By the time we got there I felt
off, but could move at least. The doctor did an EKG because of my history of
heart problems but found nothing wrong. He told me to try to avoid the person
who was causing me stress and then discharged me.</p>

<h2 id="the-inadequate-response">The (Inadequate) Response</h2>

<p>I told my adviser I would not work with the bully or in the HLT group anymore.
He agreed to this and found me another service job to work on, which was
great. Removing myself from the group stopped the abuse, but I felt like I had let
her win.</p>

<p>For his part, my advisor—whom I love and would highly recommend—did not
handle the response correctly. While he helped to get me away from my abuser,
he would later privately defend her to me, implying that she was difficult to
work with but was a really good scientist. Perhaps he was trying to explain
why she would escape punishment, but I did not need to hear that; I needed his
support. Of course, my bully suffered no consequences, and she is still
working at CERN to this day.</p>

<h2 id="the-end">The End</h2>

<p><a href="https://doi.org/10.1038/d41586-018-06040-w">Bullying in academia</a> is not something you expect when starting
your PhD program. Universities, national and international labs, and even your
research group do not have good systems in place to <a href="https://boss.blogs.nytimes.com/2012/09/26/what-do-you-do-with-the-brilliant-jerk/">deal with brilliant
jerks</a>. I was able to escape and get on with my life, but I hope that
future generations of grad students and academics won’t have to. Until then,
consider my anecdote in the same vein as my previous one: <a href="/blog/should-i-get-a-phd/"><em>Should I Get a
PhD?</em></a>; advice for aspiring academics about a topic they might not have
considered.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:tenure">
      <p>The more responsibilities you could claim, the better your chances of <a href="/blog/should-i-get-a-phd/">landing that rare tenure-track position</a> most post docs desired. <a href="#fnref:tenure" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:training">
      <p>Most of the previous on-calls were veterans who did not need in depth training or documentation. They understood the system because they had built it. <a href="#fnref:training" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
      

      

      
      
        <summary type="html"><![CDATA[Getting my PhD was mostly a great experience, but one woman made my life hell for a short time at CERN. It's tough to write about, but I thought I owed it to myself and others.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-academic-bully/the_dunce_by_harold_copping.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/my-academic-bully/the_dunce_by_harold_copping.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My Son’s Language Development</title>
      <link href="https://alexgude.com/blog/my-sons-words/" rel="alternate" type="text/html" title="My Son’s Language Development" />
      <published>2018-09-30T00:00:00-07:00</published>
      <updated>2018-09-30T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/my_sons_words</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-sons-words/"><![CDATA[<p>My son Theo was born in the summer of 2016. My wife and I knew his language
development would be interesting, because she and her parents speak Cantonese
and my parents speak Spanish (although they did not really pass it on to me).
Being huge nerds, we decided to record his progress.</p>

<h2 id="the-data">The Data</h2>

<p>We collected the data using a Google form. We decided a “word” was a phrase or
sign that Theo associated with a specific concept. This meant that when he
babbled “mama” or “baba” at 10 months we did not count those because he didn’t
have a clear association. We did count some made-up words where he had an
association, but these were mostly in sign language where he would invent
them to convey his meaning.</p>

<p>We found data collection to be difficult and error prone. Deciding when Theo
had a clear association was hard because the best indicator was that he used
the word multiple times for the same thing. Sometimes these reuses would be
separated by many days, forcing us to try to remember when he first used it.</p>

<p>As Theo got older and much better at language, we ran into new issues. First,
he was learning so fast that we had trouble keeping up and remembering if a
word had been recorded or not. Second, he became so good at mimicking sounds
that he would repeat words back to you several times, but not remember them
later.</p>

<p>Still, we think the data is a pretty good representation of his language
development. I’ll spend a future post exploring some of the quality issues.</p>

<p>You can find the Jupyter notebook used to perform this analysis
<a href="/files/my-sons-words//Theo's%20first%20words.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/my-sons-words//Theo's%20first%20words.ipynb">rendered on Github</a>). The data can be found
<a href="/files/my-sons-words//theo_words.csv">here</a>.</p>

<h2 id="development">Development</h2>

<p>Theo spoke his first word, in Cantonese, at 14 months. He had been babbling
for a long time before that, but without associating the sounds with meaning.
He then picked up four signs from <a href="https://en.wikipedia.org/wiki/Baby_sign_language">baby sign language</a> before
speaking his second word a month later. We suspect Theo picked up signing
quickly because it was a universal language in our house; Mom would only speak
Cantonese and I would only speak English, but we both responded to his signs.</p>

<p>Theo’s language development is plotted below, showing the number of words he
could speak in each “language” as a function of how old he was.</p>

<p><a href="/files/my-sons-words//theo_total_words_linear.svg"><img src="/files/my-sons-words//theo_total_words_linear.svg" alt="A plot showing the number of words my son could speak as a function of
age." /></a></p>

<p>Theo continued adding signs and Cantonese words for three months before he spoke
a word in English, the language I speak to him. That is when he also started
mimicking animal sounds. At 18 months he spoke his first word in Spanish, the
language my parents speak to him.</p>

<p>Theo slowly added words, week by week, until right before he turned 2. At 23
months his language acquisition exploded. He started the period knowing 25
Cantonese words. He knew 50 a month later, and almost 100 after two
months. Theo is now 26 months old and knows almost 200 words in Cantonese.
English exploded also, but a little later. At 25 months he knew about 40
English words, and now he knows over 100 a month and a half later!</p>

<p>His Spanish development took off at around the same time, but quickly
plateaued, only to take off again recently. The reason is simple: neither of
us speaks it well, but my parents do. The quick rise happened when he was
visiting his grandparents regularly, and the plateau is when we stopped
visiting for a few months while my parents were out of town. Now that we are
visiting them again, he has started picking up more words.</p>

<h2 id="the-words">The Words</h2>

<p>I plotted a selection of some of Theo’s first words in each language below.
Notice that I have switched to a log plot for the <em>y</em>-axis to better show the
beginnings of each language.</p>

<p><a href="/files/my-sons-words//theo_first_words.svg"><img src="/files/my-sons-words//theo_first_words.svg" alt="A plot showing the first words my son could speak as a function of
age." /></a></p>

<p>In a future post I’ll explore when Theo learned different groups of words
(colors, numbers, foods, etc.), but for now here are some of the fun words
Theo learned:</p>

<ul>
  <li>
    <p><strong>Little Brother</strong> (Cantonese): Theo has a brother, Cory, who is 18 months
younger than him; it only took Theo a few weeks to learn what the new intruder
was called.</p>
  </li>
  <li>
    <p><strong>Google</strong> (English): We have a few <a href="https://en.wikipedia.org/wiki/Google_Home">Google Homes</a> in our
apartment and so we say the trigger word, “Hey Google”, several times a day.
Theo quickly picked it up. We knew he was saying “Google” and not babbling
“Gaga” because he would either point at the device or grab it and shout at it
when saying it.</p>
  </li>
  <li>
    <p><strong>Cookie</strong> (Spanish): Theo’s first word in Spanish was “Cookie”. We can
blame the grandparents for this, as they love giving him sweets.</p>
  </li>
  <li>
    <p><strong>Monkey</strong> (Sign): “Monkey” is signed by holding one hand out with palm up
and bouncing the other on top of it, palm facing out and fingers spread. He
invented this sign after we used it while singing <a href="https://en.wikipedia.org/wiki/Five_Little_Monkeys">Five Little
Monkeys</a>.</p>
  </li>
  <li>
    <p><strong>Pig</strong> (Sign): “Pig” is signed by rubbing your chin between thumb and index
finger. He modeled this sign after the hand motions I would make while reading
<a href="https://en.wikipedia.org/wiki/The_Three_Little_Pigs">The Three Little Pigs</a>, specifically during the line “Not by the
hair on my chiny chin chin”.</p>
  </li>
</ul>

<h2 id="other-writings-on-language-development">Other Writings on Language Development</h2>

<p>If you enjoyed this article, here are all the other articles I wrote about
<a href="/topics/childhood-language/">language development</a>!</p>

<ul class="card-grid">
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/all-my-sons-words-comparison/miners_children_belva_mine_kentucky_nara.jpg" alt="Black and white photo of two young boys hanging out a window, their faces smudged with soot.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/all-my-sons-language-development-comparison/">
      <strong>Comparison of My Three Sons’ Language Development</strong>
    </a>
<br />
I recorded the words my sons spoke as they learned our various languages and now I compare how each developed! Read on to find out how each son learned.  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <img src="https://alexgude.com/files/my-third-sons-words/taylor_and_alfred_by_j_w_orr.jpg" alt="A woodcut by J. W. Orr showing a father giving his son a picture book in a richly appointed study.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-third-sons-words/">
      <strong>My Third Son’s Language Development</strong>
    </a>
<br />
We tracked my third son's language development word by word. Here, in plots, is how he learned to speak. Take a look!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <img src="https://alexgude.com/files/my-sons-words-comparison/coal_miners_child_in_grade_school_lejunior_harlan_county_kentucky.jpg" alt="Black and white photo of a young boy at a school desk.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-sons-language-development-comparison/">
      <strong>Comparison of My Two Sons’ Language Development</strong>
    </a>
<br />
Being a nerd dad, I recorded all the words my first two sons spoke as they learned them. Now, I compare their language development rate!  </div>
</li>
<li class="article-card">
  <div class="card-element card-image">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <img src="https://alexgude.com/files/my-second-sons-words/teaching_punctuation_by_j_w_orr.png" alt="A woodcut by J. W. Orr showing a man using a blackboard to teach young children punctuation.
" />
    </a>
  </div>
  <div class="card-element card-text">
    <a href="https://alexgude.com/blog/my-second-sons-words/">
      <strong>My Second Son’s Language Development</strong>
    </a>
<br />
My second son is a little over two years old. We tracked every word he's spoken to watch his language development, and now you can observe it too!  </div>
</li>
</ul>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="childhood-language" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[My son is a little over two and unfortunately he has two huge nerds for parents. We tracked every word he's spoken to watch his language development, and now you can join us!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" />
        <media:content medium="image" url="https://alexgude.com/files/my-sons-words/Articulation_by_j_w_orr.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Should I Go To Insight Data Science?</title>
      <link href="https://alexgude.com/blog/should-i-go-to-insight/" rel="alternate" type="text/html" title="Should I Go To Insight Data Science?" />
      <published>2018-08-21T00:00:00-07:00</published>
      <updated>2018-08-21T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/should_i_go_to_insight</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/should-i-go-to-insight/"><![CDATA[<p><strong>Update:</strong> Insight has <a href="https://web.archive.org/web/20240901135520/https://insightfellows.com/financing"><em>significantly changed its funding model</em></a>.
I have added a note to the section <a href="#insight-is-free-but-expensive">Insight is <em>Free</em>, but
<strong>EXPENSIVE</strong></a>.</p>

<figure>
  
  <a href="/files/insight//stanford_dish_in_the_early_morning_hours_by_brianhama_cc_by_sa_4.jpg">
    <img src="/files/insight//stanford_dish_in_the_early_morning_hours_by_brianhama_cc_by_sa_4.jpg" alt="The Stanford Dish, a large radio telescope, in the early morning
    light. The sky is purple and blue, and the dish is a white metal-lattice
    structure glowing softly pink in the light. At the top of the dish is a
    red light." decoding="async" />
  </a>
  
  
  
  <figcaption><a href="https://commons.wikimedia.org/wiki/File:Stanford_Dish_in_the_early_morning_hours.jpg"><em>Stanford Dish in the early morning hours</em></a>, © <a href="https://en.wikipedia.org/wiki/User:Brianhama">Brian Hama</a> (<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC-BY-SA 4.0</a>)</figcaption>
  
</figure>

<p>Following <a href="https://twitter.com/nerdneha">Neha’s</a> <a href="https://twitter.com/math_rachel/status/822958139343446016">advice</a> about writing a blog post for any
email that you’ve written over and over, I thought I’d answer a question I get
often: “I’m trying to get a job in data science; should I attend
<a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight</a>?”</p>

<p>I’m qualified to answer this question because during the summer of 2015 I was
an Insight Fellow in the data science track (DS-SV-2015B). That, of course, is
also my bias: I found my current and previous job through Insight, made a
bunch of friends, and am happy with the experience in general. I am also a
technical adviser, meaning I mentor Fellows weekly in exchange for a small bit
of equity in the company.<sup style="anchor-name:--fnref-disclaimer" id="fnref:disclaimer"><a href="#fn:disclaimer" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<h2 id="what-is-insight">What is Insight?</h2>

<p>Insight is a program to help people find jobs in data science related fields.
Their initial program was the <a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">data science</a> track—which I will
discuss in the remainder of this post—but they have expanded to cover
various other data-related fields as well. The data science program is offered
in multiple cities: Boston, New York, Seattle, Toronto, and the Bay Area.
There is also a remote session that is entirely online. They accept only PhDs,
mostly from technical fields but also a few other technically minded folk with
liberal arts PhDs. Each program runs for about three months and at the end you
will (ideally) have a job.</p>

<h2 id="what-would-i-do-there">What Would I Do There?</h2>

<p>Insight isn’t like graduate school. While there are a few talks covering a
variety of useful topics—SQL, startups and equity, tech interviews,
etc.—the majority of the learning is project based. You pick a project and
then spend a few weeks (feverishly) working on it, learning by doing and by
talking to the people around you. If you are interested in seeing what one of
these projects looks like, I wrote about mine in detail:
<a href="/blog/whereto-photo/">WhereTo.Photo</a>.</p>

<p>The program is structured as follows:</p>

<ul>
  <li>
    <p>4 weeks of building a data project</p>
  </li>
  <li>
    <p>3 weeks of demoing the project at various companies</p>
  </li>
  <li>
    <p>3+ weeks of interviewing with companies you demoed at</p>
  </li>
</ul>

<p>During the first few weeks, companies come by to present what they do and why
they need data scientists. After all the companies have presented you put
together a list of the companies that you are interested in. You will get to
demo at <strong>some</strong>, <strong>but not all</strong>, of them, because the number of people who
can demo at any particular company is limited.</p>

<p>Over the next few weeks you show off your project at some of the companies
that you selected (eight per Fellow during my session) in what Insight calls a
demo. The demo consists of a short presentation about the project you built, a
live demonstration of it, and a few minutes for questions. <a href="https://docs.google.com/presentation/d/1BwKT9-hDt0jHaCS6qjVlPptYKe5CkCwQUpqdT4GI8hM/edit?usp=sharing">My demo slides are
here</a> if you want to see what that looks like.</p>

<p>Some of the companies you demo at will want to take the next step and
interview you. When not demoing, you practice for those interviews and brush
up on all the topics you are unsure of with the help of the other Fellows and
mentors.</p>

<p>Throughout the program you receive guidance from a range of people:</p>

<ul>
  <li>
    <p><em>Insight staff</em>, who are often outstanding Fellows from previous sessions
hired to lead the program.</p>
  </li>
  <li>
    <p><em>Industry mentors</em>, who have worked in data science for some time, such as
myself.</p>
  </li>
  <li>
    <p><em>Previous Fellows</em> who come back to help teach the next generation, also
such as myself.</p>
  </li>
  <li>
    <p><em>The other Fellows</em> who make up your session.</p>
  </li>
</ul>

<p>This last group <strong>shouldn’t be underestimated</strong>: you learn a lot from your
peers as you ask each other for help with your projects. Everyone is an expert
at something, and a novice at other things.</p>

<h2 id="why-should-i-go">Why Should I Go?</h2>

<figure>
  
  <a href="/files/insight//stanford_dish_by_jawed_cc_by_3.jpg">
    <img src="/files/insight//stanford_dish_by_jawed_cc_by_3.jpg" alt="A landscape of the Stanford hills including the Stanford Dish
    shrouded in light fog. The sun is setting behind the dish making the sky
    appear bright white." decoding="async" />
  </a>
  
  
  
  <figcaption><a href="https://commons.wikimedia.org/wiki/File:Stanford_Dish.jpg"><em>Stanford Dish</em></a> (cropped), © <a href="https://en.wikipedia.org/wiki/User:Jawed">Jawed</a> (<a href="https://creativecommons.org/licenses/by/3.0/deed.en">CC-BY 3.0</a>)</figcaption>
  
</figure>

<h3 id="lots-of-exposure">Lots of Exposure</h3>

<p>Insight is a great way to get a quick survey of about thirty companies in the
area that are all hiring data scientists. I knew of the big name companies in
Silicon Valley (Apple, Google, Facebook, etc.) but not really any of the
medium or small companies. One of those small companies was the perfect fit
for me and I spent two great years working there; I never would have found
them without Insight to make the introduction.</p>

<h3 id="a-killer-network">A Killer Network</h3>

<p>Insight is also a great way to build a professional network. There are over
fifteen hundred alumni scattered all across the country. Friends from my
session now work at a wide variety of companies, including:
AdRoll,
Airbnb,
Facebook,
Google,
Instagram,
Intuit,
Kabaam!,
LinkedIn,
Netflix,
Salesforce,
SambaTV,
Silicon Valley Data Science,
Stitch Fix,
and
VEVO.</p>

<p>This network is incredibly valuable—it is how I found my second job—and it
would have been hard to build without Insight.</p>

<h3 id="great-friends-and-community">Great Friends and Community</h3>

<p>Finally, I’ve made great friends from among the people who went through
Insight with me. This might seem like a minor point, but if you’re picking up
your life and moving across the country it is great to have a group you can
hang out with. I regularly go on bike rides, play board games, or just get
dinner with the people from my session.</p>

<h2 id="sounds-great-whats-the-catch">Sounds Great! What’s the Catch?</h2>

<h3 id="insight-is-free-but-expensive">Insight is <em>Free</em>, but <strong>EXPENSIVE</strong></h3>

<p><strong>Update:</strong> Insight has <a href="https://web.archive.org/web/20240901135520/https://insightfellows.com/financing"><em>significantly changed its funding model</em></a>
since I wrote this article. They now require either a refundable upfront
payment, or an <a href="https://en.wikipedia.org/wiki/Income_share_agreement">income sharing agreement</a>.</p>

<p><del>Insight doesn’t charge the Fellows anything, and they even offer a limited
number of need-based scholarships to help cover expenses.</del> Still, living
without a job for three or four months is hard. Rent in the Bay Area will set
you back around $2000 a month at least. My wife and I had a good amount of
savings, but it was just about wiped out renting for three months without a
second source of income. We did not have to fall back on our credit cards, but
I know Fellows who needed to.</p>

<p>Insight is also not your employer, so you will have to get health insurance
somewhere else. My graduate school insurance lasted over the summer, but this
is not the case for everyone.</p>

<h3 id="not-everyone-finds-a-job-in-three-months">Not Everyone Finds a Job in Three Months</h3>

<p><a href="https://web.archive.org/web/20191215174644/https://www.insightdatascience.com/faq">Insight says that the vast majority of Fellows find a job in the
industry</a>, and this is true, but they don’t promise that it will happen
within the three months. Everyone in my session who stuck with it got a job
that they loved in data science, but for a few of them it took a long time.
About a third of my class had offers within a week or two, another third took
about a month or slightly longer, and the last third dragged out over a few
months with one or two of my friends taking more than four months. I’ve known
Fellows in other sessions who spent six months interviewing before receiving
an offer.</p>

<h3 id="you-need-to-know-how-to-code">You Need to Know How To Code</h3>

<p>Data scientists need to know how to work with data, but they also need to know
how to code. <a href="https://medium.com/insight-data-science/preparing-for-the-transition-to-data-science-e9194c90b42c">A good grasp of at least one language is a must</a>, and
it would be very tough to pick up coding from scratch at Insight. If you can’t
write <a href="https://imranontech.com/2007/01/24/using-fizzbuzz-to-find-developers-who-grok-coding/">FizzBuzz</a> then you will have catching up to do. This is
especially true for machine learning positions, which have interviews that
more closely follow the “standard” developer interview pattern of writing code
on the whiteboard over and over and over again.</p>

<h2 id="final-thoughts">Final Thoughts</h2>

<p>I believe in giving someone all the information they need to make an informed
decision. That’s why I wrote this post. I think Insight is worth the costs,
but I wanted to make those costs known ahead of time. In the end, I would go
through the program again. It helped my career get off to a great start by
finding the perfect first company for me. Could I have gotten a data science
job without Insight? Probably, but the experience certainly would have been
more painful, and I don’t think I would be as happy with the outcome because
the network I built has really helped expand my possibilities.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:disclaimer">
      <p>For the record, they also buy me dinner when I come in to mentor after work. <a href="#fnref:disclaimer" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[Insight promises an easy transition from academia to a career in data science or machine learning, but is it the right program for you? I have a few words of advice offered from experience.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/insight/stanford_dish_in_the_early_morning_hours_by_brianhama_cc_by_sa_4.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/insight/stanford_dish_in_the_early_morning_hours_by_brianhama_cc_by_sa_4.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Improving An Old Supernova Plot</title>
      <link href="https://alexgude.com/blog/supernova-plot-improvements/" rel="alternate" type="text/html" title="Improving An Old Supernova Plot" />
      <published>2018-07-14T00:00:00-07:00</published>
      <updated>2018-07-14T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/supernova_plot_improvements</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/supernova-plot-improvements/"><![CDATA[<p>Ten years ago, while an undergraduate, I worked as a <a href="https://en.wikipedia.org/wiki/Supernova_Cosmology_Project">supernova cosmologist</a>.
During that time, I made this plot of the peculiar <a href="https://en.wikipedia.org/wiki/SN_2002cx">supernova 2002cx</a> showing its
spectrum at four different points in time. As I <a href="/blog/matplotlib-blitting-supernova/">mentioned in my previous post on creating
animated plots (which utilized supernova data in the examples)</a>, the spectrum
tells us what is going on during the explosion and what elements are present.</p>

<p><a href="/files/supernova-plot-update//SN_2002cx_Spectra_log_old.svg"><img src="/files/supernova-plot-update//SN_2002cx_Spectra_log_old.svg" alt="The spectrum of Supernova 2002cx at four different times during the
explosion." /></a></p>

<p>It is not a bad plot—it conveys the information it is required to—but it has
a lot of room for improvement! This shouldn’t come as a surprise since it was one of the
earliest plots I made in my scientific career. Some of the shortcomings of the plot include:</p>

<ul>
  <li>
    <p>The spectra overlap a bit.</p>
  </li>
  <li>
    <p>Poor utilization of available space due to the need for large margins to accommodate the legend.</p>
  </li>
  <li>
    <p>The axes titles are wrong and do not include the units they measure.</p>
  </li>
  <li>
    <p>The tick labels collide at the corner.</p>
  </li>
</ul>

<p>The code that generated this plot can be found <a href="/files/supernova-plot-update//Old%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/supernova-plot-update//Old%20Plot.ipynb">rendered on Github</a>). It is not very good but, in my defense,
it <em>is</em> almost a decade old.</p>

<h2 id="improvements">Improvements</h2>

<p>I had not thought of the plot for years, until I ran into it again while
browsing Wikipedia. Using the experience I have gained since then to fix
and re-release it to the world seemed like the right thing to do. The
result of my improvements is below:</p>

<p><a href="/files/supernova-plot-update//SN_2002cx_Spectra_log.svg"><img src="/files/supernova-plot-update//SN_2002cx_Spectra_log.svg" alt="The same spectrum of Supernova 2002cx at four different times during the
explosion, but updated to better convey the information." /></a></p>

<p>The code that generated the improved plot can be found <a href="/files/supernova-plot-update//New%20Plot.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/supernova-plot-update//New%20Plot.ipynb">rendered on Github</a>). The code was improved as well, something
I will cover in a future post.</p>

<p>Why is this plot an improvement? It fixes the obvious errors, and reduces the
clutter and unused space so that the presented information is foremost.</p>

<p>To start, I removed the legend and replaced it with a label next to each spectrum.
This not only helps the reader to quickly identify each line, it also cuts down on the
space needed for supplemental information. Consequently, the margins can be reduced, allowing
for more of the available space to be used to display the full range of the data.</p>

<p>Then, I fixed the axis labels to indicate what they measure; for example,
adding the units to the x-axis, and relaying that the y-axis is the log of the flux with an offset.
I also removed the values on the y-axis because they are meaningless; each spectrum is area
normalized and then arbitrarily offset to prevent it from obscuring the other spectra.</p>

<p>Finally, I cleaned up a couple of things: the spectra no longer overlap, the
axes do not collide, the title and labels are larger to be more readable, and
I removed the spines on the top and right side to reduce the feeling of
clutter.</p>

<p>It is not perfect—I am still missing the units of the flux (because,
honestly, I forgot what they are)—but I think it is clearly better than the original.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[I learned to use matplotlib more than ten years ago. Around that time, I made a plot of supernova 2002cx for Wikipedia, but it was not terribly good. So this year, I updated it!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/supernova-plot-update/virgo_by_sidney_hall.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/supernova-plot-update/virgo_by_sidney_hall.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Python Patterns: max Instead of if</title>
      <link href="https://alexgude.com/blog/python-patterns-max-not-if/" rel="alternate" type="text/html" title="Python Patterns: max Instead of if" />
      <published>2018-06-14T00:00:00-07:00</published>
      <updated>2018-06-14T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/python_patterns_max_not_if</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/python-patterns-max-not-if/"><![CDATA[<p>When writing Python, I often have to look through a set of objects, determine
a score for each one of them, and save both the best score and object
associated with it. For example, looking for the highest scoring word that I
can make in <a href="https://en.wikipedia.org/wiki/Scrabble">Scrabble</a> with the letters I currently have.</p>

<p>One way to do this is to loop over all the objects and use a placeholder to
remember the best one seen so far, like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Set up placeholder variables
</span><span class="n">best_score</span> <span class="o">=</span> <span class="mi">0</span>
<span class="n">best_word</span> <span class="o">=</span> <span class="bp">None</span>

<span class="c1"># Try all possible words, saving the best one seen
</span><span class="k">for</span> <span class="n">word</span> <span class="ow">in</span> <span class="nf">possible_words</span><span class="p">(</span><span class="n">my_letters</span><span class="p">):</span>
  <span class="n">score</span> <span class="o">=</span> <span class="nf">score_word</span><span class="p">(</span><span class="n">word</span><span class="p">)</span>

  <span class="k">if</span> <span class="n">score</span> <span class="o">&gt;</span> <span class="n">best_score</span><span class="p">:</span>
    <span class="n">best_score</span> <span class="o">=</span> <span class="n">score</span>
    <span class="n">best_word</span> <span class="o">=</span> <span class="n">word</span>
</code></pre></div></div>

<p>You have probably written this logic before in some of your own code. The code
is not that complicated, but we can still improve its readability with a quick
tweak.</p>

<h2 id="simplifying-with-max">Simplifying with <code class="language-plaintext highlighter-rouge">max()</code></h2>

<p>What does <code class="language-plaintext highlighter-rouge">if score &gt; best_score</code> remind you of? The way we might implement
the <code class="language-plaintext highlighter-rouge">max()</code> function! Using <code class="language-plaintext highlighter-rouge">max()</code> helps us simplify the code nicely:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Set up placeholder variables
</span><span class="n">best_seen</span> <span class="o">=</span> <span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="bp">None</span><span class="p">)</span>

<span class="c1"># Try all possible words, saving the best one seen
</span><span class="k">for</span> <span class="n">word</span> <span class="ow">in</span> <span class="nf">possible_words</span><span class="p">(</span><span class="n">my_letters</span><span class="p">):</span>
  <span class="n">score</span> <span class="o">=</span> <span class="nf">score_word</span><span class="p">(</span><span class="n">word</span><span class="p">)</span>

  <span class="n">newest_seen</span> <span class="o">=</span> <span class="p">(</span><span class="n">score</span><span class="p">,</span> <span class="n">word</span><span class="p">)</span>
  <span class="n">best_seen</span> <span class="o">=</span> <span class="nf">max</span><span class="p">(</span><span class="n">best_seen</span><span class="p">,</span> <span class="n">newest_seen</span><span class="p">)</span>
</code></pre></div></div>

<p>Storing all the data together in a single tuple means that assignment and
comparison are now handled all at once. This makes it less likely that we will
mix up one of the assignments, and makes it clearer what we’re doing.</p>

<p>There is one potential pitfall here: <code class="language-plaintext highlighter-rouge">max()</code> picks the tuple with the largest
first element (the score in our case), which is what we want. But, if the
first elements are the same in both tuples, <code class="language-plaintext highlighter-rouge">max()</code> continues through the
remaining elements until the tie is broken. So if two words have the same
score, <code class="language-plaintext highlighter-rouge">max()</code> will then compare the words next, which it does lexically.</p>

<p>To have <code class="language-plaintext highlighter-rouge">max()</code> only compare the first element, we can use the <code class="language-plaintext highlighter-rouge">key</code>
parameter. The <code class="language-plaintext highlighter-rouge">key</code> parameter takes a function that is called on each object
and returns another object to use in the comparison. We can use it to select
just the first entry like so:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Set up placeholder variables
</span><span class="n">best_seen</span> <span class="o">=</span> <span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="bp">None</span><span class="p">)</span>

<span class="c1"># Try all possible words, saving the best one seen
</span><span class="k">for</span> <span class="n">word</span> <span class="ow">in</span> <span class="nf">possible_words</span><span class="p">(</span><span class="n">my_letters</span><span class="p">):</span>
  <span class="n">score</span> <span class="o">=</span> <span class="nf">score_word</span><span class="p">(</span><span class="n">word</span><span class="p">)</span>

  <span class="n">newest_seen</span> <span class="o">=</span> <span class="p">(</span><span class="n">score</span><span class="p">,</span> <span class="n">word</span><span class="p">)</span>
  <span class="n">best_seen</span> <span class="o">=</span> <span class="nf">max</span><span class="p">(</span><span class="n">best_seen</span><span class="p">,</span> <span class="n">newest_seen</span><span class="p">,</span> <span class="n">key</span><span class="o">=</span><span class="k">lambda</span> <span class="n">x</span><span class="p">:</span> <span class="n">x</span><span class="p">[</span><span class="mi">0</span><span class="p">])</span>
</code></pre></div></div>

<h2 id="even-simpler">Even Simpler</h2>

<p>In the above examples we wanted to save both the score and the word, but what
if we only cared about the word that generated the highest score, not the
score itself? Then there is an even simpler way!</p>

<p>By default <code class="language-plaintext highlighter-rouge">max()</code> uses the standard comparison operator, but we can change
that to use our <code class="language-plaintext highlighter-rouge">score_word()</code> using the same <code class="language-plaintext highlighter-rouge">key</code> argument from above. Then
we have:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">words</span> <span class="o">=</span> <span class="nf">possible_words</span><span class="p">(</span><span class="n">my_letters</span><span class="p">)</span>
<span class="n">best_word</span> <span class="o">=</span> <span class="nf">max</span><span class="p">(</span><span class="n">words</span><span class="p">,</span> <span class="n">key</span><span class="o">=</span><span class="n">score_word</span><span class="p">)</span>
</code></pre></div></div>

<p>Which gives us a very compact (and relatively fool proof) pattern, with all
the looping and placeholders pushed into the implementation of <code class="language-plaintext highlighter-rouge">max()</code>.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="python" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[I often have to loop over a set of objects to find the one with the greatest score. You can use an if statement and a placeholder, but there are more elegant ways!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/patterns/max_not_if.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/patterns/max_not_if.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">My PhD Thesis, In Short</title>
      <link href="https://alexgude.com/blog/my-phd-thesis/" rel="alternate" type="text/html" title="My PhD Thesis, In Short" />
      <published>2018-05-20T00:00:00-07:00</published>
      <updated>2018-05-20T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/my_phd_thesis</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/my-phd-thesis/"><![CDATA[<p>Three years ago today, I defended <a href="https://hdl.handle.net/11299/175445">my thesis</a> and graduated from the
University of Minnesota with a <a href="/blog/should-i-get-a-phd/">PhD</a> in high energy particle physics. As
part of that endeavor, I spent years traveling to and from <a href="https://cern.ch"><abbr title="European Organization for Nuclear Research">CERN</abbr></a> where I
studied the decay of <a href="https://en.wikipedia.org/wiki/W_and_Z_bosons">Z bosons</a> into <a href="https://en.wikipedia.org/wiki/Electron">electrons</a>. I wrote an esoteric
thesis on a very specific part of this decay which I have, because no one is
ever going to read it, attempted to summarize below in a more accessible
fashion.</p>

<h2 id="measurement-of-the-phistar-distribution-of-z-bosons-decaying-to-electron-pairs-with-the-cms-experiment-at-a-center-of-mass-energy-of-8-tev">Measurement of the phistar distribution of Z bosons decaying to electron pairs with the <abbr title="Compact Muon Solenoid">CMS</abbr> experiment at a center-of-mass energy of 8 TeV</h2>

<p>The <a href="https://en.wikipedia.org/wiki/Standard_Model">Standard Model</a> of particle physics is one of the most accurate
theories humans have ever come up with, describing nature almost exactly as we
observe it. Yet, for all its accuracy, we know it is still incomplete, and
even in the more complete areas there are regions where it is very difficult
to calculate what will happen. One of these regions is low energy <a href="https://en.wikipedia.org/wiki/Quantum_chromodynamics">quantum
chromodynamics</a> (<abbr title="Quantum Chromodynamics">QCD</abbr>). This is the part of the theory that handles the
interaction of <a href="https://en.wikipedia.org/wiki/Proton">protons</a>, <a href="https://en.wikipedia.org/wiki/Neutron">neutrons</a>, the <a href="https://en.wikipedia.org/wiki/Quark">quarks</a> that make them up,
and the <a href="https://en.wikipedia.org/wiki/Gluon">gluons</a> that bind them together.</p>

<p>For <a href="https://hdl.handle.net/11299/175445">my thesis</a>, I studied the interaction \(pp \to Z \to e^-e^+\),
that is, two protons colliding to produce a <a href="https://en.wikipedia.org/wiki/W_and_Z_bosons">Z boson</a> which then decays to
two <a href="https://en.wikipedia.org/wiki/Electron">electrons</a>.<sup style="anchor-name:--fnref-electron" id="fnref:electron"><a href="#fn:electron" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> This lets us explore the low energy <abbr title="Quantum Chromodynamics">QCD</abbr> region
(the \(pp\) collision) using particles (the \(e^-e^+\) pair) which do not have
any messy <abbr title="Quantum Chromodynamics">QCD</abbr> interactions to confound the data. More specifically, I looked at
the transverse momentum, \(Q_T\), of the Z boson, or how much the boson was
moving in the direction transverse to the proton beams at the time it decayed.</p>

<h3 id="the-large-hadron-collider-and-the-compact-muon-solenoid">The Large Hadron Collider and the Compact Muon Solenoid</h3>

<p>The <a href="https://en.wikipedia.org/wiki/Large_Hadron_Collider">Large Hadron Collider</a> (<abbr title="Large Hadron Collider">LHC</abbr>) is a very large, and very high energy,
particle collider on the border between Switzerland and France. It takes protons
and accelerates them along a circular track until they are traveling close to
the speed of light, and then smashes them together. This creates a region of
space with a lot of energy which is dissipated by creating new particles.</p>

<p>Looking at these collisions allows us to test the Standard Model at very high
energies where undiscovered particles may exist. Multiple smaller particle
accelerators feed into the <abbr title="Large Hadron Collider">LHC</abbr> (as shown below) and act like on-ramps, getting
the particles up-to-speed before they enter the <abbr title="Large Hadron Collider">LHC</abbr>. The <abbr title="Large Hadron Collider">LHC</abbr> hosts four major
experiments: <a href="https://en.wikipedia.org/wiki/ATLAS_experiment"><abbr title="A Toroidal LHC Apparatus">ATLAS</abbr></a>; <a href="https://en.wikipedia.org/wiki/ALICE_experiment"><abbr title="A Large Ion Collider Experiment">ALICE</abbr></a>; <a href="https://en.wikipedia.org/wiki/LHCb_experiment"><abbr title="LHC Beauty">LHCb</abbr></a>; and <abbr title="Compact Muon Solenoid">CMS</abbr>, the
<a href="https://en.wikipedia.org/wiki/Compact_Muon_Solenoid">Compact Muon Solenoid</a>, which was my experiment.</p>

<figure>
  
  <a href="/files/my-phd-thesis//alex_lhc_layout.svg">
    <img src="/files/my-phd-thesis//alex_lhc_layout.svg" alt="A diagram showing the LHC and the location of the four major
    experiments. Also shown are the smaller accelerators the feed protons into
    the LHC." decoding="async" />
  </a>
  
  
  
  <figcaption>A diagram showing the LHC and the location of the four major
    experiments. Also shown are the smaller accelerators the feed protons into
    the LHC.</figcaption>
  
</figure>

<p><abbr title="Compact Muon Solenoid">CMS</abbr> is a 14,000 ton experiment run by nearly 3000 physicists and engineers. It
is built around a point on the <abbr title="Large Hadron Collider">LHC</abbr> where protons collide and measures the
speed, mass, and flight direction of all the particles that are created after
the two protons collide. Some people like to think of it as a very large
camera, and that actually isn’t so bad an analogy: like a camera, <abbr title="Compact Muon Solenoid">CMS</abbr> measures
particles using silicon (and some other materials) and saves the output
somewhere to be viewed later.</p>

<figure>
  
  <a href="/files/my-phd-thesis//cms-color-white.png">
    <img src="/files/my-phd-thesis//cms-color-white.png" alt="A cutaway diagram of the CMS detector showing the various pieces
    that make it up." decoding="async" />
  </a>
  
  
  
  <figcaption>Cutaway model of CMS by Tai Sakuma and Thomas McCauley (CC-BY 3.0)</figcaption>
  
</figure>

<p>The cutaway above shows the detector (with silhouette for scale).<sup style="anchor-name:--fnref-sakuma" id="fnref:sakuma"><a href="#fn:sakuma" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> The
collision point is in the center. The detector is built in layers around this
collision point, with each layer designed to measure a different property of a
particle. By combining the measurements from all the layers we can tell which
particles were created and what direction they traveled in.</p>

<p>Doing research with the data generated by <abbr title="Compact Muon Solenoid">CMS</abbr> has some subtleties, but it is
essentially just counting. We predict how many events with a certain
characteristic, for example containing two high energy electrons, we should see
based on our understanding of the Standard Model. Then we count the number of
events that <abbr title="Compact Muon Solenoid">CMS</abbr> recorded with those characteristics and compare the data to our
prediction.</p>

<h3 id="transverse-momentum">Transverse Momentum</h3>

<p>I said I studied low energy <abbr title="Quantum Chromodynamics">QCD</abbr>, but then immediately introduced the highest
energy collider in the world. This is not, as it turns out, a contradiction.
While the <abbr title="Large Hadron Collider">LHC</abbr> is very high energy, all that energy is directed along the
beamline; the protons have almost no energy transverse to the
beamline,<sup style="anchor-name:--fnref-transverse" id="fnref:transverse"><a href="#fn:transverse" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> and so this gives us the low energy <abbr title="Quantum Chromodynamics">QCD</abbr> system that we
desire. The \(Z \to ee\) decay is a great way to study this low energy regime
because neither the Z nor the electrons interact via <abbr title="Quantum Chromodynamics">QCD</abbr>, so the only <abbr title="Quantum Chromodynamics">QCD</abbr>
effects in the decay chain are from the initial proton-proton collision. This
makes it an easy to measure signal.</p>

<p>I measured the transverse momentum, \(Q_T\), of the Z boson, which describes
the way the boson moves just before it decays. Measuring this not only tells us
about the low energy regime of <abbr title="Quantum Chromodynamics">QCD</abbr>, but it also helps to constrain the mass of
the W boson, which is otherwise hard to measure. The W mass is interesting
because it helps determine some fundamental quantities in the Standard Model,
and because there is some disagreement between the measured value and the
predicted value.</p>

<p>But actually, I didn’t measure \(Q_T\), I measured a new variable called
\(\phi^{*}\), which measures the same effect as \(Q_T\), but is more robust
against shortcomings in the detector design. \(\phi^{*}\) measures the angle
between the two electrons instead of their energy, which is easier to measure
accurately due to the design of <abbr title="Compact Muon Solenoid">CMS</abbr>.</p>

<h3 id="backgrounds-event-selection-and-other-issues">Backgrounds, Event Selection, and Other Issues</h3>

<p>The events I was interested in had two electrons in them, but not every event
with two electrons comes from a Z decay. Events that produce two electrons for
other reasons are called <em>background</em> events, and the primary difficulty in
any high energy experiment is separating the background from the events we are
interested in, called <em>signal</em> events. We use simple selection rules, called
<em>cuts</em>, to select the events we want. This selection happens in two stages:</p>

<ol>
  <li>
    <p>As the collisions are happening in the detector (at a rate of 40 million
times a second), a fast data processing system called the trigger applies
very simple cuts to select interesting events out of the huge number of
uninteresting ones.</p>
  </li>
  <li>
    <p>After the data has been saved to disk, I apply a more precise set of cuts
to select just the events that I think are likely to come from the \(Z \to
ee\) decay.</p>
  </li>
</ol>

<p>For example, in that second set of cuts I required two electrons, one with
very high energy and in the middle of the detector (where we get the most
accurate measurement), and one other with a little less energy anywhere in the
detector. Even with this selection some background events slip through;
figuring out how many is most of the work done in my thesis, because when we
count up events, we naturally do not want to include the background.</p>

<p>The way we try to figure out what the background looks like is through <a href="https://en.wikipedia.org/wiki/Monte_Carlo_method">Monte
Carlo experiments</a> (<abbr title="Monte Carlo">MC</abbr>) where we generate virtual data using the math
prescribed by the Standard Model. We tune this generated data on other particle
decays that look similar to, but are not, the ones we are interested in. For
example, since I cared about double electron events, we tuned our <abbr title="Monte Carlo">MC</abbr> on events
with one electron and one <a href="https://en.wikipedia.org/wiki/Muon">muon</a> (which is very much like an electron,
but heavier).</p>

<p>Below is an example of what the <abbr title="Monte Carlo">MC</abbr> predicted the signal (blue) and backgrounds
(everything else) would look like for my experiment. The black points are all
the events selected for the analysis, so you can see the agreement is not
perfect.</p>

<figure>
  
  <a href="/files/my-phd-thesis//z_peak.svg">
    <img src="/files/my-phd-thesis//z_peak.svg" alt="A plot showing the Z mass peak in data, with stacked histograms
    showing the estimated contribution from the background and signal events." decoding="async" />
  </a>
  
  
  
  <figcaption>A plot showing the Z mass peak in data, with stacked histograms
    showing the estimated contribution from the background and signal events.</figcaption>
  
</figure>

<p>There were other issues to work through as well, like estimating how good <abbr title="Compact Muon Solenoid">CMS</abbr>
was at measuring various things. These biases were mostly corrected or estimated
by using <abbr title="Monte Carlo">MC</abbr> events and putting them through the analysis pipeline. This lets us
check what the analysis would have measured, while having information about the
true underlying event.</p>

<h3 id="results">Results</h3>

<p>The results of six years of my life are shown in the (absolutely hideous, I
confess) plot below:</p>

<figure>
  
  <a href="/files/my-phd-thesis//final_result.svg">
    <img src="/files/my-phd-thesis//final_result.svg" alt="A plot showing the final result comparing the measured
    phi star distribution to the distributions predicted by various QCD Monte
    Carlo simulations." decoding="async" />
  </a>
  
  
  
  <figcaption>A plot showing the final result comparing the
    measured \(\phi^*\) distribution to the distributions predicted by various
    QCD Monte Carlo simulations.</figcaption>
  
</figure>

<p>Each collection of points in the top plot is a count of the events, for a
specific bin in \(\phi^*\). The black points are the count of events seen in
<abbr title="Compact Muon Solenoid">CMS</abbr>, and the colored points are the number of events predicted by various Monte
Carlo generator programs that use different approximations of the Standard
Model. The bottom plot is the ratio of <abbr title="Monte Carlo">MC</abbr> events over observed events. If the
generators were perfectly simulating reality, all the points would be right at
1, but of course they are not perfect so the points drift up and down. The blue
points (from the <a href="http://madgraph.physics.illinois.edu">MadGraph</a> generator) do the best job, but even they
miss by up to 5%.</p>

<h2 id="in-summary">In Summary</h2>

<p>So that’s it! I measured the angle between pairs of electrons in the <abbr title="Compact Muon Solenoid">CMS</abbr>
detector, and compared it to the angle predicted by the Standard Model. The
result can be used to fine-tune the Monte Carlo generators we use so that
future measurements have better estimates. But will my result be used for that?
It’s unclear.</p>

<p>I graduated quickly without publishing the results of my thesis in a journal.
In fact, my thesis was embargoed for six months to let my adviser and post doc
tidy it up for publication, but various complications ultimately prevented that
from happening.</p>

<p>That is the reality of scientific research though: spending years of your life
chasing down a subject cared about by only a handful of people. That is part of
the reason I wanted to write this post, in the hope that more than a handful of
people would read it, and learn about what I did.</p>

<!-- prettier-ignore -->
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:electron">
      <p>Well, an electron and a positron, but particle physicists just call those electrons because precision is for mathematicians. <a href="#fnref:electron" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:sakuma">

      <p><span class="citation">Sakuma, T, and McCauley, T. <a href="https://dx.doi.org/10.1088/1742-6596/513/2/022032">“Detector and Event Visualization with SketchUp at the CMS Experiment”</a> <cite>Journal of Physics: Conference Series</cite>. vol. 513 Track 2. 2014. doi: <a href="https://doi.org/10.1088/1742-6596/513/2/022032">10.1088/1742-6596/513/2/022032</a>.</span>
Image available under <a href="https://creativecommons.org/licenses/by/3.0/">CC-BY
3.0</a> <a href="#fnref:sakuma" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:transverse">
      <p>Space has three dimensions. The beam moves in one of those dimensions; the other two dimensions are the ones <em>transverse to the beam line</em>. <a href="#fnref:transverse" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
      

      

      
      
        <summary type="html"><![CDATA[I graduated from the University of Minnesota in June, 2015. I wrote an esoteric thesis about Z boson decay, which I explain here.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/my-phd-thesis/20110610--Building_40-,_or_-My_Home_Away_From_Home-.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/my-phd-thesis/20110610--Building_40-,_or_-My_Home_Away_From_Home-.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Making Animations Quickly with Matplotlib Blitting</title>
      <link href="https://alexgude.com/blog/matplotlib-blitting-supernova/" rel="alternate" type="text/html" title="Making Animations Quickly with Matplotlib Blitting" />
      <published>2018-04-07T00:00:00-07:00</published>
      <updated>2018-04-07T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/matplotlib_blitting_supernova</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/matplotlib-blitting-supernova/"><![CDATA[<p>Animations are a great way to show the passage of time in a plot. I have used
animation to show <a href="/blog/raspberry-pi-reboot-times/">how long my Raspberry Pis take to reboot</a> and <a href="/blog/popular-names/">how
the popularity of names changed in the US</a>.</p>

<p>But making animations in <a href="https://matplotlib.org">matplotlib</a> can take a long time. Not
just to write the code, but waiting for it to run! The easiest, but slowest,
way to make an animation is to redraw the entire plot every frame. Using this
method, it took roughly 20 minutes to render a single animation for my <a href="/blog/popular-names/">names
post</a>! Fortunately there is a significantly faster alternative:
matplotlib’s <a href="https://en.wikipedia.org/wiki/Bit_blit">animation blitting</a>. Blitting increased rendering speed by
a factor of 20!</p>

<h2 id="the-data">The Data</h2>

<p>We will plot the spectrum of <a href="https://en.wikipedia.org/wiki/SN_2011fe">Supernova 2011fe</a> from <a href="https://doi.org/10.1051/0004-6361/201221008">Pereira et
al.</a><sup style="anchor-name:--fnref-pereira_cite" id="fnref:pereira_cite"><a href="#fn:pereira_cite" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> by the <a href="https://snfactory.lbl.gov">Nearby Supernova
Factory</a>.<sup style="anchor-name:--fnref-aldering_cite" id="fnref:aldering_cite"><a href="#fn:aldering_cite" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> The <a href="https://en.wikipedia.org/wiki/Astronomical_spectroscopy">spectrum</a> of a supernova tells
us about what is going on in the explosion, so looking at a time series tells
us how the explosion is evolving.</p>

<p>The data is available <a href="https://snfactory.lbl.gov/snf/data/SNfactory_Pereira_etal_2013_SN2011fe.tar.gz">here</a>. The notebook with all the code is
<a href="/files/matplotlib-blitting-supernova/Matplotlib%20Animation%20Blitting%20Example%20-%20Supernova%20Spectra.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/matplotlib-blitting-supernova/Matplotlib%20Animation%20Blitting%20Example%20-%20Supernova%20Spectra.ipynb">rendered on Github</a>). The code in the notebooks
is complete, including doc strings and comments, while I have stripped down
the examples below for clarity.</p>

<p>This is the animation we will be making:</p>

<div class="video-gif">
  <video style="display:block; width:100%; height:auto;" autoplay="" muted="" loop="loop">
    <source src="/files/matplotlib-blitting-supernova/sn2011fe_spectral_time_series.mp4" type="video/mp4" />
  </video>
</div>

<p>It shows the amount of light (flux) the telescope saw as a function of the
wavelength of light. The data was only sampled once every few days, so to make
the animation smooth we will linearly interpolate the data. This is
implemented by the function <code class="language-plaintext highlighter-rouge">flux_from_day(day)</code>, which returns a numpy array
of flux values for a specific day. The details of how the function works can
be found in the <a href="/files/matplotlib-blitting-supernova/Matplotlib%20Animation%20Blitting%20Example%20-%20Supernova%20Spectra.ipynb">notebook</a>.</p>

<h2 id="blitting">Blitting</h2>

<p>Blitting breaks the animation into two components: the unchanging background
elements, and the <a href="https://matplotlib.org/users/artists.html">artist objects</a> that are updated each frame. It
requires us to write three functions:</p>

<ul>
  <li>
    <p><a href="#init_fig-function"><code class="language-plaintext highlighter-rouge">init_fig()</code></a>: draws the static background</p>
  </li>
  <li>
    <p><a href="#frame_iter-function"><code class="language-plaintext highlighter-rouge">frame_iter()</code></a>: yields the <code class="language-plaintext highlighter-rouge">frame_data</code> needed to draw each
update</p>
  </li>
  <li>
    <p><a href="#update_artists-function"><code class="language-plaintext highlighter-rouge">update_artists(frame_data)</code></a>: takes <code class="language-plaintext highlighter-rouge">frame_data</code> and
updates the artists</p>
  </li>
</ul>

<p>The artists that are updated each frame must be kept in an iterable container.
A normal list will work, but a more convenient way to do this is using a
<a href="https://docs.python.org/2/library/collections.html#collections.namedtuple"><code class="language-plaintext highlighter-rouge">namedtuple</code></a> (which I <a href="/blog/python-patterns-namedtuple/">discuss in detail in another
post</a>). This will let us access the different artists by name, for
example <code class="language-plaintext highlighter-rouge">artists.flux_line</code>, instead of having to remember their index number.</p>

<h3 id="init_fig-function">init_fig Function</h3>

<p>The <code class="language-plaintext highlighter-rouge">init_fig()</code> function draws the background of the animation. It takes no
arguments and must return an iterable of the artists to be updated every
frame, which in our case are contained in the namedtuple discussed above.</p>

<p>Our example function sets the labels, the title, and the range of the plot. It
is here where we would draw anything else that is unchanging, like the legend,
or some text labels, if we needed to. Here it is:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">init_fig</span><span class="p">(</span><span class="n">fig</span><span class="p">,</span> <span class="n">ax</span><span class="p">,</span> <span class="n">artists</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s">Initialize the figure, used to draw the first
    frame for the animation.
    </span><span class="sh">"""</span>
    <span class="c1"># Set the axis and plot titles
</span>    <span class="n">ax</span><span class="p">.</span><span class="nf">set_title</span><span class="p">(</span><span class="sh">"</span><span class="s">Supernova 2011fe Spectrum</span><span class="sh">"</span><span class="p">,</span> <span class="n">fontsize</span><span class="o">=</span><span class="mi">22</span><span class="p">)</span>
    <span class="n">ax</span><span class="p">.</span><span class="nf">set_xlabel</span><span class="p">(</span><span class="sh">"</span><span class="s">Wavelength [Å]</span><span class="sh">"</span><span class="p">,</span> <span class="n">fontsize</span><span class="o">=</span><span class="mi">20</span><span class="p">)</span>
    <span class="n">FLUX_LABEL</span> <span class="o">=</span> <span class="sh">"</span><span class="s">Flux [erg s$^{-1}$ cm$^{-2}$ Å$^{-1}$]</span><span class="sh">"</span>
    <span class="n">ax</span><span class="p">.</span><span class="nf">set_ylabel</span><span class="p">(</span><span class="n">FLUX_LABEL</span><span class="p">,</span> <span class="n">fontsize</span><span class="o">=</span><span class="mi">20</span><span class="p">)</span>

    <span class="c1"># Set the axis range
</span>    <span class="n">plt</span><span class="p">.</span><span class="nf">xlim</span><span class="p">(</span><span class="mi">3000</span><span class="p">,</span> <span class="mi">10000</span><span class="p">)</span>
    <span class="n">plt</span><span class="p">.</span><span class="nf">ylim</span><span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="mf">1.25e-12</span><span class="p">)</span>

    <span class="c1"># Must return the list of artists, but we use a pass
</span>    <span class="c1"># through so that they aren't created multiple times
</span>    <span class="k">return</span> <span class="n">artists</span>
</code></pre></div></div>

<p>You will notice that I said the function takes no arguments, but I gave it
three anyway. It’s hard to have no inputs (without using globals), but one
trick is to use <a href="https://en.wikipedia.org/wiki/Partial_application">partial application</a>, which I will demonstrate when
we <a href="#putting-it-all-together">put it all together</a>. The function must return the list of artists to
update, but I find it’s easier to declare those outside of the function and
then pass them in as an argument.</p>

<h3 id="frame_iter-function">frame_iter Function</h3>

<p>The <code class="language-plaintext highlighter-rouge">frame_iter()</code> function is a generator that returns the data needed to
update the artist for each frame. It yields <code class="language-plaintext highlighter-rouge">frame_data</code>, which can be any
sort of Python data type or object. This function also must take no arguments,
and so like <a href="#init_fig-function"><code class="language-plaintext highlighter-rouge">init_fig()</code></a> we will use the partial application trick
to bind the arguments.</p>

<p>Our function loops over the days relative to maximum light and returns the
flux values from that day, as well as string of the day to update the text
label.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">frame_iter</span><span class="p">(</span><span class="n">from_day</span><span class="p">,</span> <span class="n">until_day</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s">Iterate through the days of the spectra and return
    flux and day number.
    </span><span class="sh">"""</span>
    <span class="k">for</span> <span class="n">day</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="n">from_day</span><span class="p">,</span> <span class="n">until_day</span><span class="p">):</span>
        <span class="n">flux</span> <span class="o">=</span> <span class="nf">flux_from_day</span><span class="p">(</span><span class="n">day</span><span class="p">)</span>
        <span class="c1"># Yield events so the function can be looped over
</span>        <span class="nf">yield </span><span class="p">(</span><span class="n">flux</span><span class="p">,</span> <span class="sh">"</span><span class="s">Day: {day}</span><span class="sh">"</span><span class="p">.</span><span class="nf">format</span><span class="p">(</span><span class="n">day</span><span class="p">))</span>
</code></pre></div></div>

<h3 id="update_artists-function">update_artists Function</h3>

<p>Once we have <a href="#frame_iter-function"><code class="language-plaintext highlighter-rouge">frame_iter()</code></a> to generate the data for each frame,
<code class="language-plaintext highlighter-rouge">update_artists()</code> is really simple. All it has to do is:</p>

<ol>
  <li>
    <p>Unpack the <code class="language-plaintext highlighter-rouge">frames_data</code>.</p>
  </li>
  <li>
    <p>Update the plot line and the text.</p>
  </li>
</ol>

<p>For the plot line we call <code class="language-plaintext highlighter-rouge">.set_data()</code> to insert the new values; for the text
we call <code class="language-plaintext highlighter-rouge">.set_text()</code>. Our function is short:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">update_artists</span><span class="p">(</span><span class="n">frames</span><span class="p">,</span> <span class="n">artists</span><span class="p">,</span> <span class="n">lambdas</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s">Update artists with data from each frame.</span><span class="sh">"""</span>
    <span class="n">flux</span><span class="p">,</span> <span class="n">day</span> <span class="o">=</span> <span class="n">frames</span>

    <span class="n">artists</span><span class="p">.</span><span class="n">flux_line</span><span class="p">.</span><span class="nf">set_data</span><span class="p">(</span><span class="n">lambdas</span><span class="p">,</span> <span class="n">flux</span><span class="p">)</span>
    <span class="n">artists</span><span class="p">.</span><span class="n">day</span><span class="p">.</span><span class="nf">set_text</span><span class="p">(</span><span class="n">day</span><span class="p">)</span>
</code></pre></div></div>

<p>Lines and text are easy to update, but other plot objects (like histograms)
are associated with multiple artists, which makes it harder to update them.
Unfortunately, the only solution is to write a much more complicated update
function for each type.</p>

<h3 id="putting-it-all-together">Putting it all together</h3>

<p>Once we’ve written the three functions, it is pretty simple to make our
animation:</p>

<ol>
  <li>
    <p>Create the figure (<code class="language-plaintext highlighter-rouge">fig</code>) and axes (<code class="language-plaintext highlighter-rouge">ax</code>).</p>
  </li>
  <li>
    <p>Create the list of artists, in this case a line (<code class="language-plaintext highlighter-rouge">plt.plot</code>) and some text
(<code class="language-plaintext highlighter-rouge">ax.text</code>).</p>
  </li>
  <li>
    <p>Partially apply the functions by binding inputs to them with <code class="language-plaintext highlighter-rouge">partial</code>.</p>
  </li>
  <li>
    <p>Create the animation object (<code class="language-plaintext highlighter-rouge">animation.FuncAnimation</code>).</p>
  </li>
  <li>
    <p>Save the animation as an <code class="language-plaintext highlighter-rouge">.mp4</code> (<code class="language-plaintext highlighter-rouge">anim.save</code>).</p>
  </li>
</ol>

<p>Here are those steps in code:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># 1. Create the plot
</span><span class="n">fig</span><span class="p">,</span> <span class="n">ax</span> <span class="o">=</span> <span class="n">plt</span><span class="p">.</span><span class="nf">subplots</span><span class="p">(</span><span class="n">figsize</span><span class="o">=</span><span class="p">(</span><span class="mi">12</span><span class="p">,</span> <span class="mi">7</span><span class="p">))</span>

<span class="c1"># 2. Initialize the artists with empty data
</span><span class="n">Artists</span> <span class="o">=</span> <span class="nf">namedtuple</span><span class="p">(</span><span class="sh">"</span><span class="s">Artists</span><span class="sh">"</span><span class="p">,</span> <span class="p">(</span><span class="sh">"</span><span class="s">flux_line</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">day</span><span class="sh">"</span><span class="p">))</span>
<span class="n">artists</span> <span class="o">=</span> <span class="nc">Artists</span><span class="p">(</span>
    <span class="n">plt</span><span class="p">.</span><span class="nf">plot</span><span class="p">([],</span> <span class="p">[],</span> <span class="n">animated</span><span class="o">=</span><span class="bp">True</span><span class="p">)[</span><span class="mi">0</span><span class="p">],</span>
    <span class="n">ax</span><span class="p">.</span><span class="nf">text</span><span class="p">(</span><span class="n">x</span><span class="o">=</span><span class="mf">0.987</span><span class="p">,</span> <span class="n">y</span><span class="o">=</span><span class="mf">0.955</span><span class="p">,</span> <span class="n">s</span><span class="o">=</span><span class="sh">""</span><span class="p">),</span>
<span class="p">)</span>

<span class="c1"># 3. Apply the three plotting functions written above
</span><span class="n">init</span> <span class="o">=</span> <span class="nf">partial</span><span class="p">(</span><span class="n">init_fig</span><span class="p">,</span> <span class="n">fig</span><span class="o">=</span><span class="n">fig</span><span class="p">,</span> <span class="n">ax</span><span class="o">=</span><span class="n">ax</span><span class="p">,</span> <span class="n">artists</span><span class="o">=</span><span class="n">artists</span><span class="p">)</span>
<span class="n">step</span> <span class="o">=</span> <span class="nf">partial</span><span class="p">(</span><span class="n">frame_iter</span><span class="p">,</span> <span class="n">from_day</span><span class="o">=-</span><span class="mi">15</span><span class="p">,</span> <span class="n">until_day</span><span class="o">=</span><span class="mi">25</span><span class="p">)</span>
<span class="n">update</span> <span class="o">=</span> <span class="nf">partial</span><span class="p">(</span><span class="n">update_artists</span><span class="p">,</span> <span class="n">artists</span><span class="o">=</span><span class="n">artists</span><span class="p">,</span>
                 <span class="n">lambdas</span><span class="o">=</span><span class="n">np</span><span class="p">.</span><span class="nf">arange</span><span class="p">(</span><span class="mi">3298</span><span class="p">,</span> <span class="mi">9700</span><span class="p">,</span> <span class="mf">2.5</span><span class="p">))</span>

<span class="c1"># 4. Generate the animation
</span><span class="n">anim</span> <span class="o">=</span> <span class="n">animation</span><span class="p">.</span><span class="nc">FuncAnimation</span><span class="p">(</span>
    <span class="n">fig</span><span class="o">=</span><span class="n">fig</span><span class="p">,</span>
    <span class="n">func</span><span class="o">=</span><span class="n">update</span><span class="p">,</span>
    <span class="n">frames</span><span class="o">=</span><span class="n">step</span><span class="p">,</span>
    <span class="n">init_func</span><span class="o">=</span><span class="n">init</span><span class="p">,</span>
    <span class="n">save_count</span><span class="o">=</span><span class="nf">len</span><span class="p">(</span><span class="nf">list</span><span class="p">(</span><span class="nf">step</span><span class="p">())),</span>
    <span class="n">repeat_delay</span><span class="o">=</span><span class="mi">5000</span><span class="p">,</span>
<span class="p">)</span>

<span class="c1"># 5. Save the animation
</span><span class="n">anim</span><span class="p">.</span><span class="nf">save</span><span class="p">(</span>
  <span class="n">filename</span><span class="o">=</span><span class="sh">'</span><span class="s">/tmp/sn2011fe_spectral_time_series.mp4</span><span class="sh">'</span><span class="p">,</span>
  <span class="n">fps</span><span class="o">=</span><span class="mi">24</span><span class="p">,</span>
  <span class="n">extra_args</span><span class="o">=</span><span class="p">[</span><span class="sh">'</span><span class="s">-vcodec</span><span class="sh">'</span><span class="p">,</span> <span class="sh">'</span><span class="s">libx264</span><span class="sh">'</span><span class="p">],</span>
  <span class="n">dpi</span><span class="o">=</span><span class="mi">300</span><span class="p">,</span>
<span class="p">)</span>
</code></pre></div></div>

<p>The only tricky thing is the use of partial applications. Partial application
binds some (or all) of the arguments to the function and creates a new
function that takes fewer arguments. Essentially, it’s like setting a default
value for the arguments.</p>

<p>For the <code class="language-plaintext highlighter-rouge">update()</code> function above, we use partial application to fix some of
the arguments, while leaving the <code class="language-plaintext highlighter-rouge">frame</code> argument as one that still must be
supplied at call time. To create the <code class="language-plaintext highlighter-rouge">init()</code> and <code class="language-plaintext highlighter-rouge">step()</code> functions above, we
fully apply the parent functions, allowing the new functions to be called
without any inputs.</p>

<h2 id="a-little-extra">A Little Extra</h2>

<p>Of course, you can add a bit more to the plot, like the <a href="https://en.wikipedia.org/wiki/Photometric_system">photometrics
filters</a> used:</p>

<div class="video-gif">
  <video style="display:block; width:100%; height:auto;" autoplay="" muted="" loop="loop">
    <source src="/files/matplotlib-blitting-supernova/sn2011fe_spectral_time_series_extra.mp4" type="video/mp4" />
  </video>
</div>

<p>But that would have made the example even harder to follow. If you’re
interested, the notebook to generate that plot is <a href="/files/matplotlib-blitting-supernova/Matplotlib%20Animation%20Blitting%20Example%20-%20Supernova%20Spectra%20Extra.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/matplotlib-blitting-supernova/Matplotlib%20Animation%20Blitting%20Example%20-%20Supernova%20Spectra%20Extra.ipynb">rendered on Github</a>).</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:pereira_cite">

      <p><span class="citation">Pereira <abbr class="etal">et al.</abbr> <a href="https://doi.org/10.1051/0004-6361/201221008">“Spectrophotometric time series of SN 2011fe from the Nearby Supernova Factory”</a> <cite>Astronomy &amp; Astrophysics</cite>. vol. 554. 2013. pp. A27. doi: <a href="https://doi.org/10.1051/0004-6361/201221008">10.1051/0004-6361/201221008</a>.</span> <a href="#fnref:pereira_cite" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:aldering_cite">

      <p><span class="citation">Aldering <abbr class="etal">et al.</abbr> <a href="https://doi.org/10.1117/12.458107">“Overview of the Nearby Supernova Factory”</a> <cite>Proceedings Volume 4836, Survey and Other Telescope Technologies and Discoveries</cite>. 2002. doi: <a href="https://doi.org/10.1117/12.458107">10.1117/12.458107</a>.</span> <a href="#fnref:aldering_cite" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[Animating plots is great way to show how some quantity changes in time, but they can be slow to generate in matplotlib! Thankfully, blitting makes animating much faster! Learn how to here!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/matplotlib-blitting-supernova/m101.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/matplotlib-blitting-supernova/m101.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">The Typical Successful Career Starts with Rejection</title>
      <link href="https://alexgude.com/blog/a-career-starts-with-rejection/" rel="alternate" type="text/html" title="The Typical Successful Career Starts with Rejection" />
      <published>2018-03-14T00:00:00-07:00</published>
      <updated>2018-03-14T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/a_career_starts_with_rejection</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/a-career-starts-with-rejection/"><![CDATA[<p><a href="https://mrtz.org">Moritz Hardt</a> posted a <a href="https://twitter.com/mrtz/status/950493433822560257">tweet</a> the other day saying “It’s
that time of the year to keep in mind that the typical start of a successful
academic career is getting rejected from a bunch of good grad schools.” Lots
of people replied about how their own failures did not stop them. This post is
my own response.</p>

<p><img src="/files/a-career-starts-with-rejection/mrtz_tweet.png" alt="A screenshot of Moritz Hardt's Tweet saying: &quot;It's that time of the year to
keep in mind that the typical start of a successful academic career is getting
rejected from a bunch of good grad schools.&quot;" /></p>

<p>My story is a little different than Moritz’s. <a href="/blog/should-i-get-a-phd/">I do not have an academic
career</a>, I was rejected, not by a few good grad schools, but by all of
them, and yet, I am happy with where I ended up.</p>

<h2 id="to-grad-school">To Grad School</h2>

<p>I wanted to be a Physicist since I was fourteen, so I knew the path before me:
BA, PhD, professor. As I finished up my BA at <a href="https://en.wikipedia.org/wiki/University_of_California,_Berkeley">Berkeley</a>, I started
preparing for the GRE. But test anxiety caused me to study poorly, and
ultimately my test results were terrible. None of the graduate schools I
applied to in 2008 accepted me.</p>

<p>I tried again in 2009. Not the GRE, since thinking about it brought even more
anxiety than the first time, but simply applying once again to various
graduate schools. Thankfully, the <a href="https://en.wikipedia.org/wiki/University_of_Minnesota">University of Minnesota</a> (and more
specifically, the too kind <a href="https://cse.umn.edu/physics/yuichi-kubota">Yuichi Kubota</a>) looked past my test scores and
gave me a spot. I wasn’t the only person he gave a chance to, but I’m thankful
that I was one of them.</p>

<h2 id="and-onward">And Onward</h2>

<p>I graduated from grad school in 2015 and went to <a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight Data
Science</a>. From there I got a job doing machine learning at
<a href="https://www.lab41.org">Lab41</a>, where I published <a href="https://www.dropbox.com/s/q2bquqawpg8htgc/0956.pdf?dl=1">a paper</a> (<a href="https://arxiv.org/abs/1611.06962">arXiv</a>) and spent
two great years learning from my brilliant coworkers. Now, I lead a data
science team at <a href="https://www.intuit.com">Intuit</a>. Failing to get into graduate school in 2008
was crushing at the time, but it taught me that everyone falls, and sometimes
they just need a bit of a hand to get back up and on their way.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
      

      

      
      
        <summary type="html"><![CDATA[When I applied for graduate school I was rejected, not by a few good grad schools, but by all of them! But thanks to a little kindness, I was able to continue onward.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/a-career-starts-with-rejection/wheeler_auditorium_steps_in_1940.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/a-career-starts-with-rejection/wheeler_auditorium_steps_in_1940.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Rise and Fall of Popular Names</title>
      <link href="https://alexgude.com/blog/popular-names/" rel="alternate" type="text/html" title="Rise and Fall of Popular Names" />
      <published>2018-02-28T00:00:00-08:00</published>
      <updated>2018-02-28T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/popular_names</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/popular-names/"><![CDATA[<p>In the last few years I’ve named two sons, so I have been thinking about names
a lot. The only constraint my wife and I followed when picking names was that
they should <em>not</em> be too popular! We determine which names to avoid by looking
at data from the <a href="https://en.wikipedia.org/wiki/Social_Security_Administration">Social Security Administration</a>, which has kept track
of the name of every child born in the United States since 1880 (as long as at
least 5 people shared the name, for privacy reasons).</p>

<p>Most people aren’t like my wife and I: the top names are very, very popular!
But how has the popularity of the top names changed over time? I decided to
explore the trends by looking at names that were the most popular in the
United States for at least one year. The data is from the Social Security
Administration, and can be downloaded <a href="https://www.ssa.gov/oact/babynames/names.zip">here</a>. You can find the Jupyter
notebook used to perform this analysis <a href="/files/names/Most%20Popular%20Names%20Blit%20Same%20Time.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/names/Most%20Popular%20Names%20Blit%20Same%20Time.ipynb">rendered on
Github</a>). The code uses blitting to significantly speed up the
rendering, a <a href="/blog/matplotlib-blitting-supernova/">technique I cover in another post</a>.</p>

<h2 id="boys-names">Boy’s Names</h2>

<div class="video-gif">
  <video style="display:block; width:100%; height:auto;" autoplay="" muted="" loop="loop">
    <source src="/files/names/most_popular_us_boy_names.mp4" type="video/mp4" />
  </video>
</div>

<p><a href="/files/names/most_popular_us_boy_names.svg">Click here for a static version of the boy’s name plot.</a></p>

<p>The plot shows the fraction of boys born in a specific year given one of the
most popular names. When a name is the highest line on the chart, that name is
the most popular name in America during that time. We can not tell from the
plot which name is the second most popular in a given year because I have only
included names that reach the number one spot at some point. For example,
William is the second most popular name in 1920, but because it never reaches
the top spot it is not included in the plot.</p>

<p>Using the plot we can see that the most popular boy’s names are timeless; they
retain the top spot for decades, and remain popular even a century later. They
are also <a href="https://en.wikipedia.org/wiki/List_of_biblical_names">biblical</a>, with six out of the seven top names coming
from scripture.</p>

<p>John is the top boy’s name for four decades before being replaced by Robert.
Robert keeps the top spot for 16 years before James takes over, which stays at
the top for 14 years. David and Michael rise to peak together, and so the five
names share roughly equal popularity in the late 50s and early 60s.</p>

<p>While the other names decline together, Michael takes off and remains the top
name for almost four decades before being edged out by Jacob. By the time Noah
claims the most popular name title in 2013, all seven of the most popular boy’s
names have come down to about the same popularity.</p>

<h2 id="girls-names">Girl’s Names</h2>

<div class="video-gif">
  <video style="display:block; width:100%; height:auto;" autoplay="" muted="" loop="loop">
    <source src="/files/names/most_popular_us_girl_names.mp4" type="video/mp4" />
  </video>
</div>

<p><a href="/files/names/most_popular_us_girl_names.svg">Click here for a static version of the girl’s name plot.</a></p>

<p>By contrast, the most popular girl’s names seem to be driven by fads: Linda,
Lisa, Jennifer, Jessica, and Ashley all have meteoric rises and almost as
rapid falls. In fact, Lisa, Jennifer, and Ashley are so obscure before their
ascendancy that there are entire years where no girls are given those names!</p>

<p>Mary, though, has staying power, with the biblical name holding the top spot
for 67 years before Linda (with the help of a <a href="https://en.wikipedia.org/wiki/Linda_(1946_song)">hit single</a>)
unseats it in 1947. After Linda fades, Mary comes back for a decade before it
finally drops out of the most popular slot for good.</p>

<p>It is not until the late 90s that the fad trend is finally broken, with slow
rising Emily taking the top spot, to be quickly eclipsed by the trio of Emma,
Isabella, and Sophia, which are neck-and-neck through the 2000s.</p>

<h2 id="thoughts-and-observations">Thoughts and Observations</h2>

<p>Why are the most popular girl’s names driven by fads, but boy’s names are not?
I have a theory: throughout history, far more men have been famous than women.
This shows up in biblical names as well, with Mary being essentially the only
woman of importance in the Bible. This allows boy’s names to survive
generation to generation as parents look to famous men for inspiration and
use their names. Parents of girls had fewer options, and so new names were
able to fill the void each generation, leading to their quick rise and fall.</p>

<p>One pattern that is evident for both boy’s and girl’s names is the declining
relative popularity of the top names. In the 1880s, the most popular name for
each gender was given to 8% or 9% of all children born that year! Even into
the 1980s the top names were given to about 4% of children. But now the top
names are given to only 1% of children! It may very well be that the growth of
mass communication—newspapers, then film and radio, followed by television,
and ultimately the Internet—exposed American parents to an ever larger pool
of names, thus allowing diversity to win out in the end.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[The popularity of baby names rises and falls based on the tastes of each generation of parents. Are their preferences the same for boy's names as for girl's names? I plot the trends to find out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/names/swedish_children.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/names/swedish_children.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Should I Get a PhD?</title>
      <link href="https://alexgude.com/blog/should-i-get-a-phd/" rel="alternate" type="text/html" title="Should I Get a PhD?" />
      <published>2018-01-19T00:00:00-08:00</published>
      <updated>2018-01-19T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/should_i_get_a_phd</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/should-i-get-a-phd/"><![CDATA[<p>When I was fourteen, I knew that I was destined to become a physics professor,
so when I finished my BA at UC Berkeley<sup style="anchor-name:--fnref-bear" id="fnref:bear"><a href="#fn:bear" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> it never even crossed my mind
to do anything but go to graduate school. After all, one must first go to
graduate school to become a professor, and a professorship was in my future.
Ergo, graduate school (specifically, the University of Minnesota<sup style="anchor-name:--fnref-umn" id="fnref:umn"><a href="#fn:umn" class="footnote" rel="footnote" role="doc-noteref">2</a></sup>) was
the next step in my journey.</p>

<p>Seven years later, as a freshly minted PhD in particle physics, I went
straight to <a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight Data Science</a> to begin a career in Silicon
Valley. (I’ve written more about <a href="/blog/should-i-go-to-insight/">my Insight experience and advice for
prospective fellows here</a>.) The dream of becoming a
physics professor had long since been abandoned. I never even applied for a
<a href="https://en.wikipedia.org/wiki/Postdoctoral_researcher">postdoc position</a>, which would have been the next step in the
process.<sup style="anchor-name:--fnref-pd" id="fnref:pd"><a href="#fn:pd" class="footnote" rel="footnote" role="doc-noteref">3</a></sup></p>

<p>So what caused me to redirect my ambitions and energy? What was it that,
despite thoroughly enjoying my six years in Minnesota,<sup style="anchor-name:--fnref-year_off" id="fnref:year_off"><a href="#fn:year_off" class="footnote" rel="footnote" role="doc-noteref">4</a></sup> has led me
to wonder if I would still do it again were I to be given a second chance? And
what should my answer be when I am asked by students whether they should
pursue a physics PhD? These are not easy questions to answer, nor will any
particular answer be suitable for all, but if you are wondering if you should
follow in my footsteps (or wondering about the road not taken), then a review
of my experiences may reveal insights.</p>

<h2 id="you-learn-a-lot">You learn a lot</h2>

<p>I learned some highly specialized knowledge in graduate school—relativity,
quantum mechanics, and particle cross-sections. Other knowledge was extremely
useful, but completely unrelated to physics—how to work with large datasets,
how to be skeptical of my conclusions, how to communicate results, and how to
manage my time. But the most useful things I learned were about myself—to
have confidence in my abilities, that I was a worthwhile human, that I was
more than just a physicist.</p>

<p>I understand myself better having gone through grad school, and I am in many
ways a more complete person. I also developed some marketable skills, but what
I feel like I really gained was the confidence to sell the skills that I did
have.</p>

<h2 id="and-it-is-a-lot-of-fun">And it is a lot of fun</h2>

<p>Graduate school, if you find a good adviser, is also a lot of fun. You meet
hundreds of people who share the exact same passion as you, and you get to
spend six years having lunch and dinner parties with them all while riding
bikes and hiking and rock climbing and whatever else you like to do. I met
many of my best friends in graduate school and, with the exception of Insight,
it has been the best source of contacts in my professional network.</p>

<h2 id="but-there-are-no-jobs">But there are no jobs</h2>

<p>Graduate school was a slow yet constant realization that there were no <a href="https://www.nature.com/news/many-junior-scientists-need-to-take-a-hard-look-at-their-job-prospects-1.22879">jobs
in my field</a>. Going into grad school, I had naively assumed that
everyone, or nearly so, who wanted a professorship got one.
But as I watched brilliant postdocs leaving the field one after another, each
failing to get even a single offer after trying for years, I realized that I
had been wrong. The year I left with my PhD, my experiment of nearly 3000
scientists placed only around ten postdocs out of one-hundred into US-based
tenure track positions. Those are terrible odds.</p>

<h2 id="and-the-opportunity-cost-is-very-high">And the opportunity cost is very high</h2>

<p>You don’t spend your own money on a graduate degree,<sup style="anchor-name:--fnref-dont" id="fnref:dont"><a href="#fn:dont" class="footnote" rel="footnote" role="doc-noteref">5</a></sup> but it is still
<a href="https://en.wikipedia.org/wiki/Opportunity_cost">very expensive</a>. You spend six or seven years training for a job you
will not get. You learn a lot along the way, but you could certainly learn the
useful things in a much shorter amount of time.</p>

<p>In six years, if you get rid of the physics courses, exams, and teaching, you
could fit a lot of training and on-the-job experience. Most of the experience
I acquired during graduate school was not at all applicable to the job I have
now, and most hiring managers view it that way as well.</p>

<h2 id="but-wait-i-want-to-be-a-data-scientist">But wait, I want to be a data scientist</h2>

<p>It is true that many data scientists (and certainly most of the ones at the
companies I’ve worked for) have PhDs, but I think that is an artifact of how
new the position is. Teaching the useful skills in a much shorter amount of
time than required for a PhD is what the master’s degree programs in data
science, and eventually the undergraduate degrees, will attempt to do. And I
believe they will succeed after some trial and error.</p>

<p>So if your goal is to be a data scientist, ask yourself this: would you rather
have a master’s in data science, and four or five years of experience in that
industry, or would you rather be fresh out of a PhD program with no
“practical” experience in the field? Several of my friends are in the latter
position and I can tell you: it isn’t easy!</p>

<h2 id="would-i-do-it-again">Would I do it again?</h2>

<p>It’s hard to answer because I really like how my life has ended up. I would be
in a completely different place—in terms of my family, my career, and my
personality—had I not gone. And yet, for the reasons above, it is hard for
me to recommend the same path to others. If you absolutely know you are going
to be a professor,<sup style="anchor-name:--fnref-arent" id="fnref:arent"><a href="#fn:arent" class="footnote" rel="footnote" role="doc-noteref">6</a></sup> then you should pursue your goal and enter a PhD
program, but if you just want to break into data science, you should consider
other options.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:bear">
      <p>Go Bears! 🐻 <a href="#fnref:bear" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:umn">
      <p>Go Golden Gophers! <a href="https://twitter.com/goldythegopher/status/657228811751264256">🐿️</a> <a href="#fnref:umn" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:pd">

      <p>A postdoc is a position used to gain more experience before applying
for professorships. Graduates often spend five to six years doing multiple
postdocs, which are all but required to land a professorship. <a href="#fnref:pd" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:year_off">

      <p>I was forced to take a year off after undergrad; <a href="/blog/a-career-starts-with-rejection/">a story I
shared in another post</a>. <a href="#fnref:year_off" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:dont">
      <p>Or, at least, you should not! <a href="#fnref:dont" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:arent">
      <p>You probably aren’t. <a href="#fnref:arent" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="career-advice" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[In 2009, having little or no money in my purse, I thought I would go to graduate school in physics. But was it the right idea? And should you follow in my footsteps?]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/phd-should-i-go/alexis_carrel_at_the_1913_columbia_university_commencement.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/phd-should-i-go/alexis_carrel_at_the_1913_columbia_university_commencement.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Fate Dice: Statistics Testing Is Hard</title>
      <link href="https://alexgude.com/blog/blue-fate-dice-retest/" rel="alternate" type="text/html" title="Fate Dice: Statistics Testing Is Hard" />
      <published>2017-12-19T00:00:00-08:00</published>
      <updated>2017-12-19T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/blue-fate-dice-retest</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/blue-fate-dice-retest/"><![CDATA[<p><a href="/blog/fate-dice-statistics/">A few months ago</a> I dug up the data from my <a href="https://www.evilhat.com/home/fate-core/">Fate campaign</a>
and used it to test the dice we used for biases. I concluded that three of the
sets were fine, but that the fourth set, the blue dice, were significantly biased, with <a href="https://en.wikipedia.org/wiki/p-value"><em>p</em> &lt;
0.01</a>!</p>

<p>As a scientist, I want to know more than simply whether or not the dice are
biased; I also want to understand <em>how</em> they are biased. Is only one of the
dice actually bad? Are they all slightly biased, but only when combined
together is the bias significant? These questions could not be answered with
the data at hand as only the final total for each roll was recorded.
Fortunately, I still have the dice, so I decided to retest them!</p>

<p>That new test data is <a href="/files/fate-dice-statistics//blue_fate_dice_rolls.csv">here</a>. The old test data is <a href="/files/fate-dice-statistics//fate_dice_data.csv">here</a>.
You can find the Jupyter notebook used to make these calculations
<a href="/files/fate-dice-statistics//Blue%20Fate%20Dice%20Statistics.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/fate-dice-statistics//Blue%20Fate%20Dice%20Statistics.ipynb">rendered on Github</a>).</p>

<h2 id="individual-tests">Individual Tests</h2>

<p>To perform the test I rolled each 500 times in a row and recorded
the results, for a total of 2000 rolls. See my <a href="/blog/fate-dice-statistics/#fate-dice">previous post</a>
for a review of how Fate dice work.</p>

<p>One way to visualize all the rolls is to sum up the results for each die roll
by roll. This gives a cumulative total that wanders up and down as the die
rolls high and low results. Each die can then be modeled as a <a href="https://en.wikipedia.org/wiki/Random_walk">1-dimensional
random walk</a>. I’ve plotted the contours that 95% and 99% of
random walks lie within.</p>

<p><a href="/files/fate-dice-statistics//blue_fate_dice_cumulative_rolls.svg"><img src="/files/fate-dice-statistics//blue_fate_dice_cumulative_rolls.svg" alt="The cumulative roll values for each of the four Fate dice." /></a></p>

<p>None of the dice wander too far out of the contours, but that doesn’t
guarantee they are unbiased. For example, a die that <em>always</em> rolled 0 would
be highly biased but also stay within the contours. It is still possible that
together the dice are biased.</p>

<h2 id="group-test">Group Test</h2>

<p>Each die was rolled individually, but Fate dice are rolled four at a time and
summed. In order to mimic this with the data I generated, I took the first
roll of each of the four dice and added them together, treating that as one
roll. I repeated this process for the rest of the data to get 500 rolls for
the set of dice.</p>

<p>That gives the following distribution, where the points indicate the number
of rolls of the dice that came up with a certain value, and the grey area is
the range in which we would expect to find a result produced by a fair set of
dice 99% of the time. I discussed in detail how these regions are computed in
<a href="/blog/fate-dice-intervals/">a previous post</a>.</p>

<p><a href="/files/fate-dice-statistics//blue_fate_dice_rolls.svg"><img src="/files/fate-dice-statistics//blue_fate_dice_rolls.svg" alt="The results of the second set of blue dice rolls." /></a></p>

<p>Let me stop for a second: <em><strong>This is surprising!</strong></em></p>

<p>My <a href="/blog/fate-dice-statistics/#significance">previous test</a> shows that the blue dice were biased
at the <em>p</em> &lt; 0.01 level and yet not a single count is outside the 99% range
this time! Using a <a href="https://en.wikipedia.org/wiki/Pearson%27s_chi-squared_test">chi-squared test</a> test on the new data gives <em>p</em> =
0.66, which does not rule out the unbiased hypothesis! In fact, this new test
agrees the best with the unbiased hypothesis of all the tests <a href="/blog/fate-dice-statistics/#significance">performed last
time</a>!</p>

<p>We can compare the two tests using the same cumulative plot shown above, but
this time taking the total of all four dice as a single step.</p>

<p><a href="/files/fate-dice-statistics//blue_fate_dice_cumulative_rolls_old.svg"><img src="/files/fate-dice-statistics//blue_fate_dice_cumulative_rolls_old.svg" alt="Cumulative rolls values for the blue dice comparing the first and second tests." /></a></p>

<p>The first test, from my <a href="/blog/fate-dice-statistics/">previous post</a>, very quickly wanders
outside the 99% contour and spends much of its time there. The second test
stays solidly within the contours.</p>

<h2 id="explanation">Explanation</h2>

<p>So what explains a significant result in the first test and not in the second?
There are a few possibilities, all of which fall into two categories:
statistics and systematics.</p>

<h3 id="statistics">Statistics</h3>

<p>It is possible that the dice are biased (or fine) and the test that says
otherwise is just a statistical fluke. At <em>p</em> &lt; 0.01 that happens 1 in 100
times. Performing further tests would answer this question: biased dice would
have more results with a low <em>p</em>-value, while unbiased dice would have few.</p>

<h3 id="systematics">Systematics</h3>

<p>It is also possible, and I think more likely, that one of the tests was
performed in a biased manner. The second test was very carefully done, but the
first test was less controlled: we wrote down results when we remembered, the
person writing the results changed from day to day, and the person rolling
also changed. Further, we often remembered to start recording only <strong>after</strong> a
particularly bad roll. Performing multiple tests and looking at the
distribution of the <em>p</em> value might offer a clue indicating whether the first
test was systematically off, but it is hard to disentangle from statistical
uncertainty.</p>

<h2 id="conclusion">Conclusion</h2>

<p>So what was it, statistics or systematic? If I had to bet, I’d say that the
first test was performed poorly and that the dice are probably fine. Am I
going to test them <em><strong>again</strong></em> to check? Maybe… You will see it here if I
do!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="fun-and-games" />
        
      

      

      
      
        <summary type="html"><![CDATA[A few months ago I tested my Fate dice for biases. Now, I retest the "biased" set and see if it really is unlucky! Unfortunately, things aren't so clear...]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/fate-dice-statistics/blue_fate_dice.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/fate-dice-statistics/blue_fate_dice.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">How Fast Does a Raspberry Pi Reboot?</title>
      <link href="https://alexgude.com/blog/raspberry-pi-reboot-times/" rel="alternate" type="text/html" title="How Fast Does a Raspberry Pi Reboot?" />
      <published>2017-11-13T00:00:00-08:00</published>
      <updated>2017-11-13T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/raspberry_pi_reboot_times</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/raspberry-pi-reboot-times/"><![CDATA[<p>I own two <a href="https://en.wikipedia.org/wiki/Raspberry_Pi">Raspberry Pis</a> which are currently taped to my kitchen
cabinets. They perform a range of tasks that require an always-on, but
low-power, computer. The first one, named <a href="https://twitter.com/RaspberryPion">Raspberry Pion</a>,<sup style="anchor-name:--fnref-humor" id="fnref:humor"><a href="#fn:humor" class="footnote" rel="footnote" role="doc-noteref">1</a></sup>
seeds open-source torrents 24/7. It is a slightly older (and hence slower)
<a href="https://www.raspberrypi.org/products/raspberry-pi-2-model-b/">Raspberry Pi 2 Model B</a>. The second one, named <a href="https://twitter.com/RaspberryKaon">Raspberry
Kaon</a>,<sup style="anchor-name:--fnref-kaon" id="fnref:kaon"><a href="#fn:kaon" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> runs a VPN that I connect to when using insecure wireless
networks away from home. It is a newer <a href="https://www.raspberrypi.org/products/raspberry-pi-3-model-b/">Raspberry Pi 3 Model B</a>.</p>

<p>Both computers run <a href="https://ubuntu-mate.org/raspberry-pi/">Ubuntu Mate 16.04 for the Raspberry Pi</a>, and both
suffer from a memory leak I have not been able to track down. My solution is
to reboot the computers at 0100 every night using a <a href="https://en.wikipedia.org/wiki/Cron">cronjob</a>. They
report their status to Twitter when they come back online, which lets me know
that they have successfully rebooted and how long it took. <a href="https://twitter.com/charles_uno">One of my
friends</a> noticed that Raspberry Pion seemed to take a few seconds
longer than Raspberry Kaon, which prompted me to take a look.</p>

<p>You can find the notebook <a href="/files/raspberry-pi/Rasperry%20Pi%20Reboot%20Times.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/raspberry-pi/Rasperry%20Pi%20Reboot%20Times.ipynb">rendered on Github</a>).
Pion’s tweet data is <a href="/files/raspberry-pi/pion_tweets.csv">here</a>, and Kaon’s tweet data is
<a href="/files/raspberry-pi/kaon_tweets.csv">here</a>.</p>

<h2 id="reboot-times">Reboot Times</h2>

<p>The Raspberry Pis report the time they come back online to Twitter, <a href="https://twitter.com/RaspberryKaon/status/929272644498624513">as
follows</a>:</p>

<p><img src="/files/raspberry-pi/20171111_reboot_tweet.png" alt="An example of the status Tweet sent by Raspberry Kaon" /></p>

<p>The clocks on the Raspberry Pis are kept synchronized with a central server
using <a href="https://en.wikipedia.org/wiki/Network_Time_Protocol">NTP</a>. The network latency of sending the tweet is not an issue as
the timestamps are generated locally before being sent to Twitter. However,
the two machines differ in more than just their hardware: the Raspberry Pi 2
serves torrents meaning it has hundreds of network connections open which
might slow down its shutdown process. So this is not a completely fair
benchmark.</p>

<p>I pulled down all the reboot announcement tweets from my two Raspberry Pis and
computed the time difference in seconds from 0100. I discarded any difference
over five minutes, as these were primarily cases where the Raspberry Pi
rebooted at some other time of the day. From these I created an animated
histogram comparing the reboot times of the two computers over the 10 months
they have been running. Each month is about one second of animation.</p>

<div class="video-gif">
  <video style="display:block; width:100%; height:auto;" autoplay="" muted="" loop="loop">
    <source src="/files/raspberry-pi/raspberry_pi_reboot_times_2_vs_3_animation.mp4" type="video/mp4" />
  </video>
</div>

<p><a href="/files/raspberry-pi/raspberry_pi_reboot_times_2_vs_3.svg">Here is a static image of the plot</a> if you prefer.</p>

<p>The Raspberry Pi 2 and 3 reboot with median times of 32 and 29 seconds
respectively. The peaks are quite sharp indicating that the time to reboot is
pretty consistent. There is a peak of fast reboots for the Pi 2 at
about 25 seconds, which seem to come in clumps, which I have no good
explanation for. You can also see when I move apartments in late August; I had
no internet for awhile and so the Raspberry Pis were not plugged in, leading
to a half a second of no updates!</p>

<h2 id="animated-plots">Animated Plots</h2>

<p>A final note about animated plots: I am a huge fan of using animation to
represent the time axis because I think it makes the display of information
more intuitive. Using <code class="language-plaintext highlighter-rouge">FuncAnimation</code> from <code class="language-plaintext highlighter-rouge">matplotlib</code> was a bit tough (and I
think my code is far from optimal), but once I got it working it was a lot
faster than rendering the individual frames and creating the video afterwards.
For tips on making the animation render even faster, see my <a href="/blog/matplotlib-blitting-supernova/">post on using
blitting in <code class="language-plaintext highlighter-rouge">matplotlib</code></a>.</p>

<p>In the future I hope to make more fun animations!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:humor">
      <p>From <a href="https://en.wikipedia.org/wiki/Pion">pion</a>; physicists have notoriously terrible senses of humor. <a href="#fnref:humor" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:kaon">
      <p>After the <a href="https://en.wikipedia.org/wiki/Kaon">kaon</a>, of course. If I get another, I’ll have to name it something like <a href="https://en.wikipedia.org/wiki/J/psi_meson">J/ψ</a>, and that just doesn’t have the same ring to it. <a href="#fnref:kaon" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[My Raspberry Pis have to reboot every evening to avoid a memory leak. As they say, when you have a memory leak, make animated plots to see how fast they reboot!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/raspberry-pi/raspberry_pi_2_b_by_evan-amos.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/raspberry-pi/raspberry_pi_2_b_by_evan-amos.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Updated Caltrain Visual Schedule</title>
      <link href="https://alexgude.com/blog/caltrain-visual-schedule-updates/" rel="alternate" type="text/html" title="Updated Caltrain Visual Schedule" />
      <published>2017-10-07T00:00:00-07:00</published>
      <updated>2017-10-07T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/caltrain_visual_schedule_updates</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/caltrain-visual-schedule-updates/"><![CDATA[<p>On July 15, 2017, Caltrain changed their weekend schedule in order to allow
construction related to the <a href="https://en.wikipedia.org/wiki/Electrification_of_Caltrain">Peninsula Corridor Electrification
Project</a>. Instead of running hourly, trains now run roughly every 90
minutes—a fact I discovered when I showed up at the <a href="https://en.wikipedia.org/wiki/San_Antonio_station_(Caltrain)">San Antonio
station</a> on the 15th and learned it would be another 20 minutes before my
train would arrive. This can make it a little frustrating when trying to get
to the city to have <a href="https://knowyourmeme.com/memes/avocado-toast">avocado toast</a> with your friends.</p>

<p>A new schedule is not all bad though; it means a chance to reuse the script I
developed to produce <a href="/blog/caltrain-visual-schedule/">visual schedules for Caltrain</a>. You can find
the notebook <a href="/files/caltrain-schedule/20170715-Caltrain%20Marey%20Schedule.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/caltrain-schedule/20170715-Caltrain%20Marey%20Schedule.ipynb">rendered on Github</a>). The schedule
data is from <a href="http://www.caltrain.com/developer.html">Caltrain’s developer site</a>.</p>

<h1 id="saturday">Saturday</h1>

<p>The top schedule is the new July 15th one, the old schedule is below. Click to
enlarge.</p>

<p><a href="/files/caltrain-schedule/caltrain_saturday_20170715.svg"><img src="/files/caltrain-schedule/caltrain_saturday_20170715.svg" alt="Marey visual train schedule for caltrain on Saturday after the July 15 change" /></a>
<a href="/files/caltrain-schedule/caltrain_saturday_20170308.svg"><img src="/files/caltrain-schedule/caltrain_saturday_20170308.svg" alt="Marey visual train schedule for caltrain on Saturday before the July 15 change" /></a></p>

<p>The frequency of the trains has decreased to every 90 minutes, and the pair of
express trains now spend more time in San Francisco before departing again.
There is also an interesting pair of northbound trains that leave very close
together. These are required to keep the number of trains heading up the
peninsula the same as the number heading down.</p>

<h1 id="sunday">Sunday</h1>

<p>Again, the top schedule is the new one, the bottom one is the old schedule.</p>

<p><a href="/files/caltrain-schedule/caltrain_sunday_20170715.svg"><img src="/files/caltrain-schedule/caltrain_sunday_20170715.svg" alt="Marey visual train schedule for caltrain on Sunday after the July 15 change" /></a>
<a href="/files/caltrain-schedule/caltrain_sunday_20170308.svg"><img src="/files/caltrain-schedule/caltrain_sunday_20170308.svg" alt="Marey visual train schedule for caltrain on Sunday before the July 15 change" /></a></p>

<p>The new Sunday schedule is identical to the new Saturday schedule, except that
the first and last northbound trains have been removed, along with the final
two southbound trains. Interestingly, the Sunday trains actually run a bit
later under the new schedule, with the last train leaving Diridon at just past
2200, instead of at 2100.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[In July, Caltrain updated their weekend schedule to allow time to do track work, so I updated my Marey/Ibry/Serjev visual schedules to see how it changed!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/caltrain-schedule/an_outbound_sp_commuter_train_sequence_by_roger_puta.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/caltrain-schedule/an_outbound_sp_commuter_train_sequence_by_roger_puta.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Tech Interviews: Respect Everyone’s Time</title>
      <link href="https://alexgude.com/blog/interviews-respect-time/" rel="alternate" type="text/html" title="Tech Interviews: Respect Everyone’s Time" />
      <published>2017-09-18T00:00:00-07:00</published>
      <updated>2017-09-18T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/interviews-respect-time</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/interviews-respect-time/"><![CDATA[<p>A lot has been written about how to perform technical interviews. What
questions to ask, what questions not to ask, why the questions we used to ask
are bad. I have nothing of value to add to that debate.<sup style="anchor-name:--fnref-haoyi" id="fnref:haoyi"><a href="#fn:haoyi" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>What I do have to add is one guiding principle for the interviewing company:
<strong>Do not waste anyone’s time!</strong></p>

<p>Do not waste the time of the candidate by making them jump through endless
hoops. Do not waste the time of your own people (the ones performing the
interview) by losing candidates due to poor communication. To this end, I have
two simple pieces of advice that very few of the companies I have interviewed
with have followed.</p>

<h2 id="structure-just-enough">Structure: Just Enough</h2>

<p>A few companies I have interviewed with have had labyrinthine processes such
as:</p>

<ul>
  <li>
    <p>multiple phone screens</p>
  </li>
  <li>
    <p>multiple days of on-site interview loops</p>
  </li>
  <li>
    <p>phone screens as well as take-home projects</p>
  </li>
</ul>

<p>Instead of utilizing such protracted methods, ask what you hope to learn with
a second interview (or phone screen, or data challenge) that you failed to
learn the first time, and then figure out how to learn that key piece of
information in just one round.</p>

<p>I think the following steps are enough:</p>

<ul>
  <li>
    <p>Resume screen (performed without the candidate)</p>
  </li>
  <li>
    <p>Casual chat with the candidate about the position</p>
  </li>
  <li>
    <p>Technical screen</p>
  </li>
  <li>
    <p>On-site interview loop</p>
  </li>
</ul>

<h2 id="communicate-often">Communicate: Often</h2>

<p>The vast majority of companies follow poor communication practices during the
hiring process. Often, there are long periods of silence followed by
unpredictable bursts of noise. Worse, some companies simply stop communicating
altogether! The companies that offered the best experience had recruiters who
stayed in constant contact, and who set appointments and kept them. The
absolute best even scheduled the post on-site interview loop feedback (and
hiring decision) phone call at the same time they gave me the schedule for the
on-site itself. I very much appreciated knowing that I would know the outcome
at a set time in the future.</p>

<p>Improving communication during the hiring process is simple. Communicate
before each step in the hiring process and let the candidate know what to
expect. Then communicate after each step and let the candidate know whether or
not they are continuing in the hiring process. Ideally, schedule these
interactions ahead of time so there is no doubt as to when they will happen.
This gives the candidate a fixed date on which they will have an answer, and
gives a deadline to the hiring team to make a decision.</p>

<h2 id="thats-it">That’s It</h2>

<p>It is tough to hire, but it is not tough to be better at structuring
interviews. Respect the time of all parties involved and you will be well on
your way to a better experience for the candidates and the interviewers.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:haoyi">
      <p>If that is what you’d like though, I think <a href="https://www.lihaoyi.com/post/HowtoconductagoodProgrammingInterview.html">Li Haoyi’s <em>How to conduct a good Programming Interview</em></a> is pretty good. <a href="#fnref:haoyi" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="interviewing" />
        
          <category term="opinions" />
        
      

      

      
      
        <summary type="html"><![CDATA[Interviewing is notoriously agonizing for both the candidate and the company! But it could be much better! Here I propose one guiding principle to make it easier on everyone.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/interviews/watch_and_suite_jeremy_beadle.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/interviews/watch_and_suite_jeremy_beadle.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Fate Dice Intervals</title>
      <link href="https://alexgude.com/blog/fate-dice-intervals/" rel="alternate" type="text/html" title="Fate Dice Intervals" />
      <published>2017-08-14T00:00:00-07:00</published>
      <updated>2017-08-14T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/fate_dice_intervals</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/fate-dice-intervals/"><![CDATA[<p>Last month <a href="/blog/fate-dice-statistics/">I checked my Fate dice for biases</a>. One of the things I
did was plot an interval for the 4dF outcomes (-4 through 4) we expect from a
fair set of dice 99% of the time. In this post I will look at four different
methods of computing those regions. While writing this post, I came upon a
<a href="https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm">NIST handbook page</a> that covers the same topic; check it out too!</p>

<p>As per usual, you can find the Jupyter notebook used to perform these
calculations and make the plots <a href="/files/fate-dice-statistics//Fate%20Dice%20Expectation%20Regions.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/fate-dice-statistics//Fate%20Dice%20Expectation%20Regions.ipynb">rendered on
Github</a>).</p>

<h2 id="normal-approximation">Normal Approximation</h2>

<p>One of the simplest ways of determining how often we expect an outcome to
appear is to assume that the distribution of results is
<a href="https://en.wikipedia.org/wiki/Normal_distribution">Gaussian</a>.<sup style="anchor-name:--fnref-clt" id="fnref:clt"><a href="#fn:clt" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> If the outcome has a probability <em>P</em>, and the dice
are thrown <em>N</em> times, then the range of expected results is:</p>

<!-- prettier-ignore-start -->

\[M_{\pm} = NP \pm z \sqrt{NP(1-P)}\]

<!-- prettier-ignore-end -->

<p>Where <a href="https://en.wikipedia.org/wiki/Standard_score"><em>z</em></a> is the correct value for the interval (2.58 for 99%) and the
two M values are the lower (for minus) and upper (for plus) bounds on the
region.</p>

<p>Using this approximation yields values that are close to exact, with the
exception that they allow negative counts for rare outcomes. The values (the
negative outcomes -4 through -1 are removed, because the distribution is symmetric)
are:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Outcome</th>
      <th style="text-align: right">Lower Bound</th>
      <th style="text-align: right">Upper Bound</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">97.47</td>
      <td style="text-align: right">147.42</td>
    </tr>
    <tr>
      <td style="text-align: left">1</td>
      <td style="text-align: right">79.64</td>
      <td style="text-align: right">126.58</td>
    </tr>
    <tr>
      <td style="text-align: left">2</td>
      <td style="text-align: right">45.05</td>
      <td style="text-align: right">83.84</td>
    </tr>
    <tr>
      <td style="text-align: left">3</td>
      <td style="text-align: right">13.01</td>
      <td style="text-align: right">38.55</td>
    </tr>
    <tr>
      <td style="text-align: left">4</td>
      <td style="text-align: right">-0.06</td>
      <td style="text-align: right">12.95</td>
    </tr>
  </tbody>
</table>

<h2 id="wilson-score-interval">Wilson Score Interval</h2>

<p>The <a href="https://en.wikipedia.org/wiki/Binomial_proportion_confidence_interval#Wilson_score_interval">Wilson score interval</a> gives a better result than the normal
approximation, but at the expense of a slightly more complicated formula.
Unlike the normal approximation, the Wilson interval is asymmetric and can not
go below 0. It is defined as:</p>

<!-- prettier-ignore-start -->

\[M_{\pm} = \frac{N^2}{N+z^2} \left[ P + \frac{z^2}{2N} \pm z \sqrt{ \frac{P \left(1 - P\right)}{N}  + \frac{z^2}{4N^2}} \,\right]\]

<!-- prettier-ignore-end -->

<p>Plugging in the numbers yields:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Outcome</th>
      <th style="text-align: right">Lower Bound</th>
      <th style="text-align: right">Upper Bound</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">99.34</td>
      <td style="text-align: right">149.02</td>
    </tr>
    <tr>
      <td style="text-align: left">1</td>
      <td style="text-align: right">81.73</td>
      <td style="text-align: right">128.46</td>
    </tr>
    <tr>
      <td style="text-align: left">2</td>
      <td style="text-align: right">47.52</td>
      <td style="text-align: right">86.31</td>
    </tr>
    <tr>
      <td style="text-align: left">3</td>
      <td style="text-align: right">15.72</td>
      <td style="text-align: right">41.74</td>
    </tr>
    <tr>
      <td style="text-align: left">4</td>
      <td style="text-align: right">2.43</td>
      <td style="text-align: right">16.84</td>
    </tr>
  </tbody>
</table>

<h2 id="monte-carlo-simulation">Monte Carlo Simulation</h2>

<p>The previous two methods were quick to calculate, but only returned
approximate results. One way to determine the exact intervals is to <a href="https://en.wikipedia.org/wiki/Monte_Carlo_method">simulate
rolling the dice</a>. This is often easy to implement, but is slow due to the
high trial count required.</p>

<p>The following code (which can be found in the <a href="https://github.com/agude/agude.github.io/blob/master/files/fate-dice-statistics//Fate%20Dice%20Expectation%20Regions.ipynb">notebook</a>) will
“roll” 4dF <em>N</em> times per trial, and perform 10,000 trials:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">simulate_rolls</span><span class="p">(</span><span class="n">n</span><span class="p">,</span> <span class="n">trials</span><span class="o">=</span><span class="mi">10000</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s"> Simulate rolling 4dF N times and calculate the expectation
    intervals.
    </span><span class="sh">"""</span>

    <span class="c1"># The possible values we can select, the weights for each,
</span>    <span class="c1"># and a histogram binning to let us count them quickly
</span>    <span class="n">values</span> <span class="o">=</span> <span class="p">[</span><span class="o">-</span><span class="mi">4</span><span class="p">,</span> <span class="o">-</span><span class="mi">3</span><span class="p">,</span> <span class="o">-</span><span class="mi">2</span><span class="p">,</span> <span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="mi">0</span><span class="p">,</span> <span class="mi">1</span><span class="p">,</span> <span class="mi">2</span><span class="p">,</span> <span class="mi">3</span><span class="p">,</span> <span class="mi">4</span><span class="p">]</span>
    <span class="n">bins</span> <span class="o">=</span> <span class="p">[</span><span class="o">-</span><span class="mf">4.5</span><span class="p">,</span> <span class="o">-</span><span class="mf">3.5</span><span class="p">,</span> <span class="o">-</span><span class="mf">2.5</span><span class="p">,</span> <span class="o">-</span><span class="mf">1.5</span><span class="p">,</span> <span class="o">-</span><span class="mf">0.5</span><span class="p">,</span> <span class="mf">0.5</span><span class="p">,</span> <span class="mf">1.5</span><span class="p">,</span> <span class="mf">2.5</span><span class="p">,</span> <span class="mf">3.5</span><span class="p">,</span> <span class="mf">4.5</span><span class="p">]</span>
    <span class="n">weights</span> <span class="o">=</span> <span class="p">[</span><span class="mi">1</span><span class="p">,</span> <span class="mi">4</span><span class="p">,</span> <span class="mi">10</span><span class="p">,</span> <span class="mi">16</span><span class="p">,</span> <span class="mi">19</span><span class="p">,</span> <span class="mi">16</span><span class="p">,</span> <span class="mi">10</span><span class="p">,</span> <span class="mi">4</span><span class="p">,</span> <span class="mi">1</span><span class="p">]</span>

    <span class="n">results</span> <span class="o">=</span> <span class="p">[[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[]]</span>

    <span class="c1"># Perform a trial
</span>    <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="n">trials</span><span class="p">):</span>
        <span class="c1"># We select all n rolls "at once" using a weighted choice function
</span>        <span class="n">rolls</span> <span class="o">=</span> <span class="nf">choices</span><span class="p">(</span><span class="n">values</span><span class="p">,</span> <span class="n">weights</span><span class="o">=</span><span class="n">weights</span><span class="p">,</span> <span class="n">k</span><span class="o">=</span><span class="n">n</span><span class="p">)</span>
        <span class="n">counts</span> <span class="o">=</span> <span class="n">np</span><span class="p">.</span><span class="nf">histogram</span><span class="p">(</span><span class="n">rolls</span><span class="p">,</span> <span class="n">bins</span><span class="o">=</span><span class="n">bins</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span>

        <span class="c1"># Add the results to the global result
</span>        <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">count</span> <span class="ow">in</span> <span class="nf">enumerate</span><span class="p">(</span><span class="n">counts</span><span class="p">):</span>
            <span class="n">results</span><span class="p">[</span><span class="n">i</span><span class="p">].</span><span class="nf">append</span><span class="p">(</span><span class="n">count</span><span class="p">)</span>

    <span class="k">return</span> <span class="n">results</span>
</code></pre></div></div>

<p>After generating the trials, the intervals are computed by looking at the 0.5
percentile and the 99.5 percentile for each possible 4dF outcomes. The results
are:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Outcome</th>
      <th style="text-align: right">Lower Bound</th>
      <th style="text-align: right">Upper Bound</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">98</td>
      <td style="text-align: right">148</td>
    </tr>
    <tr>
      <td style="text-align: left">1</td>
      <td style="text-align: right">80</td>
      <td style="text-align: right">127</td>
    </tr>
    <tr>
      <td style="text-align: left">2</td>
      <td style="text-align: right">46</td>
      <td style="text-align: right">84</td>
    </tr>
    <tr>
      <td style="text-align: left">3</td>
      <td style="text-align: right">14</td>
      <td style="text-align: right">39</td>
    </tr>
    <tr>
      <td style="text-align: left">4</td>
      <td style="text-align: right">1</td>
      <td style="text-align: right">14</td>
    </tr>
  </tbody>
</table>

<h2 id="binomial-probability">Binomial Probability</h2>

<p>Simulating the rolls is guaranteed to produce the right result, but it takes a
lot of time to run. For a simple case like rolling dice, we can calculate the
intervals exactly using a little knowledge of probability. This is the method
I used in my <a href="/blog/fate-dice-statistics/">previous post</a> because it is <em>fast <strong>and</strong> exact</em>.</p>

<p>The interval indicates the expected results a fair set of dice would roll 99%
of the time, but that is exactly what probability gives as well! The
interval is therefore just the set of rolls that make up 99% of the cumulative
probability, centered around the most likely value for each outcome.
Equivalently, we can find the set of rolls that make up the very unlikely 1%,
which will come (approximately equally) from both tails of the
distribution. That is, integrate from the low side (which is a sum, since
the bins are discrete) until the cumulative probability is 0.5%, and then
repeat for the high side. The stopping points are the correct lower and upper
bounds.</p>

<p>Here is an example image showing this process for the probability distribution
of the number of zeroes rolled if the dice are thrown 522 times. The red parts
of the histogram are the results of the two integrals, each containing about
0.5% of the probability, and the grey lines mark the lower and upper bounds at
98 and 148.</p>

<p><a href="/files/fate-dice-statistics//fate_dice_probabilities.svg"><img src="/files/fate-dice-statistics//fate_dice_probabilities.svg" alt="The probability of rolling zero on 4dF a set number of times given 522 rolls." /></a></p>

<p>Each of the bins in the plot has probability given by:</p>

<!-- prettier-ignore-start -->

\[\binom{N}{M} P^M (1-P)^{N-M}\]

<!-- prettier-ignore-end -->

<p>Where <em>P</em> is the probability of rolling the outcome on one roll and <em>M</em> is the
number of time the outcome happens in <em>N</em> throws.</p>

<p>Applying this process to all the possible outcomes gives the following
results:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Outcome</th>
      <th style="text-align: right">Lower Bound</th>
      <th style="text-align: right">Upper Bound</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">98</td>
      <td style="text-align: right">148</td>
    </tr>
    <tr>
      <td style="text-align: left">1</td>
      <td style="text-align: right">80</td>
      <td style="text-align: right">127</td>
    </tr>
    <tr>
      <td style="text-align: left">2</td>
      <td style="text-align: right">46</td>
      <td style="text-align: right">84</td>
    </tr>
    <tr>
      <td style="text-align: left">3</td>
      <td style="text-align: right">14</td>
      <td style="text-align: right">39</td>
    </tr>
    <tr>
      <td style="text-align: left">4</td>
      <td style="text-align: right">1</td>
      <td style="text-align: right">14</td>
    </tr>
  </tbody>
</table>

<h2 id="comparison">Comparison</h2>

<p>Tables are great when you want exact numbers, but it is much easier to compare
the various methods using a plot. The following plot shows the predictions
from each of the four methods for the outcomes 0 through 4. The negative
outcomes (-4 through -1) are omitted because the distributions are symmetric.</p>

<p><a href="/files/fate-dice-statistics//fate_dice_regions.svg"><img src="/files/fate-dice-statistics//fate_dice_regions.svg" alt="The four different methods of computing the expected intervals." /></a></p>

<p>The Monte Carlo method and the estimate using the binomial probability agree
exactly, as expected. The naive variance method agrees well for the first few
values, but begins predicting lower intervals as the value increases, finally
ending with a nonsense negative count. The Wilson interval is consistently
higher than the other values, and this discrepancy increases as the value of
the roll increases.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:clt">
      <p>The <a href="https://en.wikipedia.org/wiki/Central_limit_theorem">central limit theorem</a> can be used to justify this approximation, but as you can see in the plot in the <a href="#binomial-probability">binomial section</a>, even for large <em>N</em> the distribution is not a perfect Gaussian. <a href="#fnref:clt" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="fun-and-games" />
        
      

      

      
      
        <summary type="html"><![CDATA[What does a "normal" distribution of rolls from a fair set of Fate dice look like? There are a lot of ways to estimate it. In this post I'll go through four methods.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/fate-dice-statistics/alphonse-mucha-fate-1920.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/fate-dice-statistics/alphonse-mucha-fate-1920.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Matching Cars with Siamese Networks</title>
      <link href="https://alexgude.com/blog/lab41-matching-cars-with-siamese-networks/" rel="alternate" type="text/html" title="Matching Cars with Siamese Networks" />
      <published>2017-08-09T00:00:00-07:00</published>
      <updated>2017-08-09T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_matching_cars_with_siamese_networks</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-matching-cars-with-siamese-networks/"><![CDATA[<p>Lab41 just finished <a href="https://github.com/Lab41/pelops"><strong>Pelops</strong></a>, a <em>vehicle re-identification
project</em> using data from fixed video cameras. <a href="/blog/lab41-object-localization-without-deep-learning/">Last time I talked about
“chipping”</a>, that is extracting an image of a vehicle from a frame
of video automatically. We found that background subtraction worked OK based
on the small amount of labeled data we had.</p>

<p>In this post I’ll go over the rest of the pipeline: <strong>feature extraction</strong> and
<strong>vehicle matching</strong>.</p>

<h2 id="feature-extraction">Feature Extraction</h2>

<p>Machine learning algorithms operate on a vector of numbers. An image can be
thought of as a vector of numbers—three numbers to define the color of each
pixel—but it turns out that taking these numbers and transforming them gives
a more <a href="https://en.wikipedia.org/wiki/Visual_descriptor">useful representation</a>. This step of taking an image and creating
a vector of useful numbers is called feature extraction. We performed feature
extraction using several different algorithms.</p>

<p>Our first method of feature extraction was an old one: Histogram of Oriented
Gradients or HOG. HOG was first proposed in the 80s, but has since found uses
in identifying pedestrians and was used by our sister lab <a href="https://medium.com/the-downlinq">CosmiQ
Works</a> to <a href="https://medium.com/the-downlinq/histogram-of-oriented-gradients-hog-heading-classification-a92d1cf5b3cc">identify boat headings</a>. HOG effectively counts the
number and direction of edges it finds in an image, and as such is useful for
finding specific objects. HOG lacks color information, so in addition to the
output from HOG, we appended a histogram of each color channel.</p>

<p>Our second method of feature extraction was based on deep learning. We took
<a href="https://arxiv.org/abs/1512.03385">ResNet50</a><sup style="anchor-name:--fnref-he" id="fnref:he"><a href="#fn:he" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> trained on <a href="https://image-net.org/">ImageNet</a>,<sup style="anchor-name:--fnref-deng" id="fnref:deng"><a href="#fn:deng" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> removed the fully
connected layers, and treated the 2048-dimension output of the convolutional
layers as our feature vector. <a href="https://cs231n.github.io/transfer-learning/">It is well known that networks trained on
ImageNet, despite being exceptionally good at identifying dogs and cats, are
also very good for general image problems</a>. It turns out the edges,
shapes, and colors learned for dogs are also, in different configurations,
useful for cars. For more details on the ResNet architecture, see my reading
group blog post.</p>

<p>Our third method of feature extraction was a <a href="https://cs231n.github.io/transfer-learning/">fine-tuned</a> ResNet50.
Pretrained networks are good at general image tasks, but they can be
“fine-tuned” to perform better on specific tasks. For Pelops that specific
task was make, model, and color identification of cars in a labeled dataset.
It is hoped that making the model better at make, model, and color detection
will generate features that are more useful for matching cars. This makes
intuitive sense: any human matching cars would use make, model, and color as
primary features.</p>

<h2 id="matching">Matching</h2>

<p>Once all of the vehicles have feature vectors associated with them, those
vectors can be used to match vehicles to each other. There are a few ways to
do this, starting with the simplest, which is to calculate a distance between
the vectors. This works great if the feature extractors are designed to make
the distance meaningful, but this is not generally the case. Neural networks
can have cost functions based on distance, but ResNet50 does not. So this
method, while attractive in simplicity, is not a good solution.</p>

<p>The second possible way of matching is to train a traditional (that is,
non-deep learning based) classifier. We trained a logistic regression model, a
random forest, and a support vector machine on top of each of the types of
feature vectors. Each model was given two feature vectors and asked to
classify them as coming from the same vehicle, or not. The training data was
balanced so that there were as many positive as negative examples. The best
accuracy these models achieved was 80%, although most struggled to pass 70%.
Accuracy is the number of true results divided by the number of total items
tested.</p>

<p>The third method of matching was to use a neural network as a classifier. Once
we added a deep learning classifier on top of our deep learning feature
extractor, we had a Siamese neural network. For details about how we trained
such a network, and for an overview of its architecture, see our blog post
here. The Siamese network performs the feature extraction and matching in one
step, and so allows optimizing both portions at the same time. This
arrangement achieved the best results by far, hitting nearly 93% accuracy on
our test set.</p>

<figure>
  
  <a href=" /files/siamese-networks//siamese_network.png ">
    <img src=" /files/siamese-networks//siamese_network.png " alt="A cartoon drawing of our Siamese network." decoding="async" />
  </a>
  
  
  
  <figcaption>A cartoon of our Siamese network architecture. The two
  convolutional blocks (CNN) output vectors which are joined together and then
  passed through a set of fully connected (FC) layers for classification.</figcaption>
  
</figure>

<h2 id="results">Results</h2>

<h3 id="dataset">Dataset</h3>

<p>In order to determine how well our various feature extraction and matching
algorithms did, we needed a labeled dataset. We used the <a href="https://ieeexplore.ieee.org/document/7553002/">VeRi
dataset</a>,<sup style="anchor-name:--fnref-veri" id="fnref:veri"><a href="#fn:veri" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> which contains pictures of 776 uniquely identified
vehicles. There are multiple pictures of each vehicle taken from 20 different
traffic cameras in China. An example of two VeRi images from Liu <em>et
al.</em><sup style="anchor-name:--fnref-liu" id="fnref:liu"><a href="#fn:liu" class="footnote" rel="footnote" role="doc-noteref">4</a></sup> is shown below.</p>

<figure>
  
  <a href=" /files/siamese-networks//trucks.jpg ">
    <img src=" /files/siamese-networks//trucks.jpg " alt="Two images from the dataset showing a front and rear view of the
  same truck." decoding="async" />
  </a>
  
  
  
  <figcaption>Two example images from VeRi showing the same truck passing two
  different cameras.</figcaption>
  
</figure>

<p>This dataset allowed us to test our performance on essentially the exact task
we were hoping to solve: re-identifying the same vehicle if it passed another
camera.</p>

<h3 id="metric">Metric</h3>

<p>The final metric we used is a cumulative matching curve (CMC). A CMC is
constructed as follows: 10 vehicles are selected in one set, and 10 in
another. These two sets have one car in common that is the correct match. The
algorithms then rank all 100 pairwise comparisons by confidence that they are
the same vehicles. The rank of the correct pair on this list of 100 pairs is
recorded. This trial is repeated for many randomly selected sets. The curve is
generated by recording what fraction of trials have the correct pair ranked at
a certain position or better.</p>

<figure>
  
  <a href=" /files/siamese-networks//cmc_plot.png ">
    <img src=" /files/siamese-networks//cmc_plot.png " alt="A plot of the effectiveness of our various methods." decoding="async" />
  </a>
  
  
  
  <figcaption>Comparison of three matching methods with a random baseline using a cumulative matching curve.</figcaption>
  
</figure>

<p>The CMC plot shows the final Siamese network compared to HOG + Color Histogram
using euclidean distance, ResNet50 with euclidean distance, and a purely
random selection of matches. The sharp rise in the Siamese CMC is because it
is very good at matching on color, so all matches where the cars share the
same color appear near the top of the rankings. The slow rise after about rank
10 is due to cases where color was not very helpful in making the match,
either because the car was a very common color, or because it was a color
easily mistaken for another (for example, yellow and white are hard for the
network to tell apart).</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:he">

      <p><span class="citation">He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian. “Deep Residual Learning for Image Recognition” <cite>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</cite>. 2016. pp. 770–778. doi: <a href="https://doi.org/10.1109/CVPR.2016.90">10.1109/CVPR.2016.90</a>.</span> <a href="#fnref:he" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:deng">

      <p><span class="citation">Deng, J. and Dong, W. and Socher, R. and Li, L.-J. and Li, K. and Fei-Fei, L. <a href="https://www.image-net.org/static_files/papers/imagenet_cvpr09.pdf">“ImageNet: A Large-Scale Hierarchical Image Database”</a> <cite>CVPR09</cite>. 2009.</span> <a href="#fnref:deng" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:veri">

      <p><span class="citation">Liu, Xinchen and Liu, Wu and Ma, Huadong and Fu, Huiyuan. “Large-scale vehicle re-identification in urban surveillance videos” <cite>2016 IEEE International Conference on Multimedia and Expo (ICME)</cite>. 2016. pp. 1–6. doi: <a href="https://doi.org/10.1109/ICME.2016.7553002">10.1109/ICME.2016.7553002</a>.</span> <a href="#fnref:veri" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:liu">

      <p><span class="citation">Liu, Xinchen and Liu, Wu and Mei, Tao and Ma, Huadong. “A Deep Learning-Based Approach to Progressive Vehicle Re-identification for Urban Surveillance” <cite>European Conference on Computer Vision</cite>. Edited by Leibe, Bastian and Matas, Jiri and Sebe, Nicu and Welling, Max. Springer International Publishing. 2016. pp. 869–884. doi: <a href="https://doi.org/10.1007/978-3-319-46475-6_53">10.1007/978-3-319-46475-6_53</a>.</span> <a href="#fnref:liu" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="lab41" />
        
          <category term="pelops" />
        
      

      

      
      
        <summary type="html"><![CDATA[Matching the same object across separate images is tough, but Siamese networks can learn to do it pretty well! Read on for details.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/siamese-networks/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/siamese-networks/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Object Localization without Deep Learning</title>
      <link href="https://alexgude.com/blog/lab41-object-localization-without-deep-learning/" rel="alternate" type="text/html" title="Object Localization without Deep Learning" />
      <published>2017-08-07T00:00:00-07:00</published>
      <updated>2017-08-07T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_object_localization_without_deep_learning</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-object-localization-without-deep-learning/"><![CDATA[<p>Lab41 is just wrapping up a vehicle re-identification project named
<a href="https://github.com/Lab41/pelops"><strong>Pelops</strong></a>. The goal was to be able to identify if the same car
drove past a fixed set of video cameras without reading the license plate.</p>

<p>We broke the project down into three parts:</p>

<ul>
  <li>
    <p><strong>Chipping</strong>: localizing vehicles within an image or frame from a video and
extract an image (chip) containing the vehicle</p>
  </li>
  <li>
    <p><strong>Feature extraction</strong>: producing a compact representation of a vehicle’s
chip suitable for machine learning</p>
  </li>
  <li>
    <p><strong>Matching</strong>: grouping chips of the same vehicle based on their feature
representations</p>
  </li>
</ul>

<p>Today I’ll show you our approach to chipping; <a href="/blog/lab41-matching-cars-with-siamese-networks/">feature extraction and matching
are covered in a later post</a>.</p>

<h2 id="chipping">Chipping</h2>

<p>There has been a lot of work done on using deep learning for object
localization in the past few years. Deep learning based methods currently
achieve state of the art results for many localization problems. In fact, our
sister lab, <a href="https://medium.com/the-downlinq">CosmiQ Works</a>, explored using these techniques, and even
developed a modified version of YOLO called <a href="https://medium.com/the-downlinq/you-only-look-twice-multi-scale-object-detection-in-satellite-imagery-with-convolutional-neural-38dad1cf7571">You Only Look Twice (YOLT)</a>
to do ship and plane localization in satellite images.</p>

<p>Deep learning seems like the best option for finding cars as well! But we
don’t use deep learning, for two reasons:</p>

<ol>
  <li>
    <p>An “off-the-shelf” network, <a href="https://arxiv.org/abs/1512.02325">Single Shot MultiBox Detector (SSD)</a><span class="nowrap">,<sup style="anchor-name:--fnref-liu" id="fnref:liu"><a href="#fn:liu" class="footnote" rel="footnote" role="doc-noteref">1</a></sup><span> performed poorly.</span></span></p>
  </li>
  <li>
    <p>We did not have much labeled data, so retraining wasn’t an option.</p>
  </li>
</ol>

<p>We were able to hand label about 200 frames of the traffic camera data in
order to test our algorithms, but did not have enough time (or, critically,
patience) to label enough vehicles to train or fine-tune a deep learning
model. Therefore, we chose an old standard in computer vision: <a href="https://en.wikipedia.org/wiki/Background_subtraction">background
subtraction</a>.</p>

<h3 id="background-subtraction">Background Subtraction</h3>

<p>What background subtraction tries to do, at its simplest, is to classify every
pixel in an image as either background or foreground. Background pixels are
ignored while foreground pixels are taken to be part of an object of interest.
To do this classification the algorithms develop background models. We used
two different algorithms with the only difference being the method in which
they model the background pixels.</p>

<p>The first method used a very simple background model and was mainly intended
as a benchmark. For each pixel it takes the median luminosity of the last ten
frames. This gives a crude estimation of what objects are “permanent” and
which are transient.</p>

<p>The second method was based on <a href="https://doi.org/10.1016/j.patrec.2005.11.005">a paper by Zivkovic and van der
Heijden</a>.<sup style="anchor-name:--fnref-zivkovic" id="fnref:zivkovic"><a href="#fn:zivkovic" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> It uses a Gaussian mixture model to estimate
the distribution of background pixels in RGB space. It learns the
distributions based on the previous frames and updates as new frames are
processed. We used the implementation in OpenCV called MOG2,
<a href="http://docs.opencv.org/3.1.0/d7/d7b/classcv_1_1BackgroundSubtractorMOG2.html"><code class="language-plaintext highlighter-rouge">cv::BackgroundSubtractorMOG2</code></a>, which also includes a model for shadow
rejection. Shadows are otherwise a tough problem to solve because shadows are
real changes in the image, but are generally not ones that are interesting.</p>

<p>After the background models are computed, the two algorithms perform the same
steps. They subtract their background model from the image leaving black
pixels where there was no change and pixels with some luminosity in areas
where there was a moving object. We then run a Gaussian blur over the image in
order to join incorrectly separated regions (e.g., where the model splits an
object in half). After the blur is applied, all pixels below a threshold value
are set to black and bounding boxes are drawn around the remaining pixels. For
both of our models, the size of the Gaussian kernel and the threshold value
are free parameters which we optimize over. The best values for these
parameters are discussed in the results section.</p>

<h3 id="data-and-labeling">Data and Labeling</h3>

<p>The data we used to test our chipping algorithms was from two sets of traffic
cameras. The first set was from San Antonio, Texas, and the second set was
from New Orleans, Louisiana. These two datasets cover a wide range of vehicle
sizes in terms of number of pixels. The Texas data is low resolution and the
cameras are mounted high above the highway giving very small chips, with the
largest being roughly 50 pixels on a side. The Louisiana data is taken with
higher resolution cameras that are mounted closer to the roadway and so the
chips are much larger. Some of the chips are up to 300 pixels on a side. An
example of the Texas data and Louisiana data follow:</p>

<figure>
  
  <a href=" /files/object-localization//freeways.png ">
    <img src=" /files/object-localization//freeways.png " alt="An example of two images from our dataset." decoding="async" />
  </a>
  
  
  
  <figcaption>An example of the Texas traffic camera data (Left) and Louisiana traffic camera data (Right). The axes indicate pixel count.</figcaption>
  
</figure>

<p>We hand labeled the images by drawing bounding boxes around the vehicles in a
series of frames. This process was tedious and so we were only able to label
about 200 frames and just under 1000 vehicles. The labeled data covered two
cameras from Texas and one from Louisiana. All the frames were captured during
the day time and good weather.</p>

<h3 id="metric">Metric</h3>

<p>We use two metrics to assess the quality of the chipping. The first is average
intersect over union (IOU) and the second is the box-wise average F1 score.
These are similar (but not identical) to the metric CosmiQ Works uses for
<a href="https://medium.com/the-downlinq/the-spacenet-metric-612183cc2ddb">SpaceNet</a>.</p>

<p>The average IOU is computed as follows: each predicted box is matched to a
ground truth box such that the combination has the highest IOU score, and no
other combination with that truth box has a higher score. Predicted boxes
without a matching truth box are given a score of 0, and truth boxes without a
matched predicted box are also given a score of 0. These IOU scores are
computed across all frames for a particular camera and the mean is computed.</p>

<p>The average F1 score uses the IOU scores computed for each matched pair above,
but with a threshold assigned. Pairs that have a higher IOU than the threshold
count as “good matches” while all others count as bad matches. Truth boxes
without matched predicted boxes are false negatives, and predicted boxes
without matched truth boxes are false positives.</p>

<p>Mean IOU gives a more fine-grained metric where high overlap is better, low
overlap is worse, and partial overlap is scored depending on how much is
matched. The F1 score, on the other hand is binary. Overlaps are either above
threshold and good, or below threshold and bad.</p>

<h3 id="results">Results</h3>

<p>The MOG2 algorithm provided by OpenCV performed best on the three different
cameras we tested, two from Texas and one from Louisiana. For the Texas data,
the optimal kernel size was around 10 pixels (although only odd sizes are
allowed), while it was nearly 100 pixels for the Louisiana data. This puts the
kernels at roughly the same scale as the size of the objects we are trying to
detect, which makes sense as smaller kernels will not suppress the noise
enough while larger ones will merge nearby objects together.</p>

<p>The optimal pixel intensity threshold value was also different between the two
datasets: about 10 for Texas and about 50 for Louisiana (where 0 is black, and
255 is white). Unlike the kernel size, we do not have a good explanation for
this.</p>

<p>The top three results for each camera in the test set are shown below, along
with example images. In the example images the truth boxes are marked in green
and the predicted boxes are in red.</p>

<h4 id="data-for-texas-ih-37-at-9th-street-camera">Data for Texas IH-37 at 9th Street Camera</h4>

<figure>
  
  <a href=" /files/object-localization//texas_9th.png ">
    <img src=" /files/object-localization//texas_9th.png " alt="An example from the Texas IH-37 at 9th Street Camera." decoding="async" />
  </a>
  
  
  
  <figcaption>Example results from the Texas traffic camera at IH-37 and 9th
  Street. The red boxes are the predicted regions, and the green boxes are the
  hand-labeled truth.</figcaption>
  
</figure>

<p>The results for the three best chipping methods for the Texas IH-37 at 9th
Street camera:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Method</th>
      <th style="text-align: right">Kernel Size</th>
      <th style="text-align: right">Threshold</th>
      <th style="text-align: right">F1 Score</th>
      <th style="text-align: right">IOU</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">11 x 11</td>
      <td style="text-align: right">8</td>
      <td style="text-align: right">0.623</td>
      <td style="text-align: right">0.273</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">13 x 13</td>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0.622</td>
      <td style="text-align: right">0.280</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">11 x 11</td>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0.614</td>
      <td style="text-align: right">0.276</td>
    </tr>
  </tbody>
</table>

<h4 id="data-for-texas-ih-37-at-jones-avenue-camera">Data for Texas IH-37 at Jones Avenue Camera</h4>

<figure>
  
  <a href=" /files/object-localization//texas_jones.png ">
    <img src=" /files/object-localization//texas_jones.png " alt="An example from the Texas IH-37 at Jones Avenue Camera." decoding="async" />
  </a>
  
  
  
  <figcaption>Example results from the Texas traffic camera at IH-37 and Jones
  Avenue. The red boxes are the predicted regions, and the green boxes are the
  hand-labeled truth.</figcaption>
  
</figure>

<p>The results for the three best chipping methods for the Texas IH-37 at Jones
Avenue camera:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Method</th>
      <th style="text-align: right">Kernel Size</th>
      <th style="text-align: right">Threshold</th>
      <th style="text-align: right">F1 Score</th>
      <th style="text-align: right">IOU</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">5 x 5</td>
      <td style="text-align: right">4</td>
      <td style="text-align: right">0.698</td>
      <td style="text-align: right">0.261</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">15 x 15</td>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0.681</td>
      <td style="text-align: right">0.277</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">11 x 11</td>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0.675</td>
      <td style="text-align: right">0.270</td>
    </tr>
  </tbody>
</table>

<h4 id="data-for-the-louisiana-traffic-camera">Data for the Louisiana Traffic Camera</h4>

<figure>
  
  <a href=" /files/object-localization//louisiana.png ">
    <img src=" /files/object-localization//louisiana.png " alt="An example from the Louisiana traffic camera." decoding="async" />
  </a>
  
  
  
  <figcaption>Example results from the Louisiana traffic camera. The red boxes
  are the predicted regions, and the green boxes are the hand-labeled truth.</figcaption>
  
</figure>

<p>The results for the three best chipping methods for the Louisiana traffic
camera:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Method</th>
      <th style="text-align: right">Kernel Size</th>
      <th style="text-align: right">Threshold</th>
      <th style="text-align: right">F1 Score</th>
      <th style="text-align: right">IOU</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">99 x 99</td>
      <td style="text-align: right">50</td>
      <td style="text-align: right">0.783</td>
      <td style="text-align: right">0.419</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">101 x 101</td>
      <td style="text-align: right">50</td>
      <td style="text-align: right">0.783</td>
      <td style="text-align: right">0.419</td>
    </tr>
    <tr>
      <td style="text-align: left">OpenCV</td>
      <td style="text-align: right">99 x 97</td>
      <td style="text-align: right">50</td>
      <td style="text-align: right">0.783</td>
      <td style="text-align: right">0.410</td>
    </tr>
  </tbody>
</table>

<h3 id="conclusion">Conclusion</h3>

<p>The results are adequate, but not great. There are a few major issues with the
MOG2 algorithm:</p>

<ul>
  <li>
    <p>Shadows still present a major problem, as seen in the image from the camera
positioned at IH-37 at Jones Ave.</p>
  </li>
  <li>
    <p>Getting the correct size for the bounding box is tough, as seen in the
Louisiana frame.</p>
  </li>
  <li>
    <p>Nearby vehicles are often merged together.</p>
  </li>
</ul>

<p>Despite these issues, background subtraction based localization is good enough
for our purpose, and easy to implement with a small amount of data. In the
future, we would focus on labeling more data so that training a deep learning
system was feasible, and focus on classifying pixels within the chips as
vehicle or background to help the down-stream algorithms focus on the
important parts of the chip.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:liu">

      <p><span class="citation">Liu, Wei and Anguelov, Dragomir and Erhan, Dumitru and Szegedy, Christian and Reed, Scott and Fu, Cheng-Yang and Berg, Alexander C. “SSD: Single Shot MultiBox Detector” <cite>Computer Vision – ECCV 2016</cite>. Edited by Leibe, Bastian and Matas, Jiri and Sebe, Nicu and Welling, Max. Springer International Publishing. 2016. pp. 21–37. doi: <a href="https://doi.org/10.1007/978-3-319-46448-0_2">10.1007/978-3-319-46448-0_2</a>.</span> <a href="#fnref:liu" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:zivkovic">

      <p><span class="citation">Zivkovic, Z. and van der Heijden, F. “Efficient adaptive density estimation per image pixel for the task of background subtraction” <cite>Pattern Recognition Letters</cite>. vol. 27. Elsevier B.V. January 6, 2006. pp. 773–780. doi: <a href="https://doi.org/10.1016/j.patrec.2005.11.005">10.1016/j.patrec.2005.11.005</a>.</span> <a href="#fnref:zivkovic" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="lab41" />
        
          <category term="pelops" />
        
      

      

      
      
        <summary type="html"><![CDATA[Finding objects in images can be hard if you have only a little data. In this post I examine a few approaches that work with few training examples!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/object-localization/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/object-localization/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Fate Dice Statistics</title>
      <link href="https://alexgude.com/blog/fate-dice-statistics/" rel="alternate" type="text/html" title="Fate Dice Statistics" />
      <published>2017-07-28T00:00:00-07:00</published>
      <updated>2017-07-28T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/fate_dice_statistics</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/fate-dice-statistics/"><![CDATA[<p>My friends and I played <a href="https://www.evilhat.com/home/fate-core/">Fate</a>, a <a href="https://en.wikipedia.org/wiki/Tabletop_role-playing_game">role-playing game</a>, for a few
years during graduate school. Over that time we developed superstitions about
the various dice we rolled. Since we were (are) huge nerds we decided to
record (almost) all of the rolls to determine if the dice really were biased.
We cursorily looked at the data when we finished playing, but I thought it
would be interesting to dig it back out and analyze it more deeply.</p>

<h2 id="fate-dice">Fate Dice</h2>

<p><a href="https://en.wikipedia.org/wiki/Fudge_(role-playing_game_system)#Fudge_dice">Fate dice</a> (also called Fudge dice) have six sides and three values
with equal probability of appearing: plus, blank, and minus. These
respectively represent +1, 0, and -1 . Four dice are rolled at a time and
their results are summed, giving a range of -4 to 4. The
<a href="https://en.wikipedia.org/wiki/Dice_notation">notation</a> for this type of roll is 4dF.</p>

<p>Figuring out the probability of rolling a value is just simple combinatorics.
These probabilities are:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Value</th>
      <th style="text-align: right">Probability</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">0</td>
      <td style="text-align: right">19/81</td>
    </tr>
    <tr>
      <td style="text-align: left">1 xor -1</td>
      <td style="text-align: right">16/81</td>
    </tr>
    <tr>
      <td style="text-align: left">2 xor -2</td>
      <td style="text-align: right">10/81</td>
    </tr>
    <tr>
      <td style="text-align: left">3 xor -3</td>
      <td style="text-align: right">4/81</td>
    </tr>
    <tr>
      <td style="text-align: left">4 xor -4</td>
      <td style="text-align: right">1/81</td>
    </tr>
  </tbody>
</table>

<h2 id="rolls">Rolls</h2>

<p>We had four sets of Fate dice, colored blue, red, black, and white. We wrote
down only the sum of each roll, since the individual dice in the set are
indistinguishable. This means that if one of the dice is biased, it will take
longer to show up than if we had been able to explore the results
individually. As per usual, you can find the Jupyter notebook used to make
these calculations <a href="/files/fate-dice-statistics//Fate%20Dice%20Statistics.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/fate-dice-statistics//Fate%20Dice%20Statistics.ipynb">rendered on Github</a>). The data
is <a href="/files/fate-dice-statistics//fate_dice_data.csv">here</a>.</p>

<p>Here are the distributions of rolls for each of the four sets of dice. The
points indicate the number of rolls that came up with a certain value, while
the grey area is the range in which we would expect to find a result produced
by a fair set of dice 99% of the time. I discuss how these regions are
computed in detail in <a href="/blog/fate-dice-intervals/">another post</a>.</p>

<p><a href="/files/fate-dice-statistics//fate_dice_rolls.svg"><img src="/files/fate-dice-statistics//fate_dice_rolls.svg" alt="The results of our rolls during out Fate campaign." /></a></p>

<p>The blue dice were rolled the most (because we thought the red and black sets
were unlucky), but visual inspection suggests that they were actually biased!
Contrary to our superstitions, the “cursed” red and black dice seem to have
been fine. The white dice have one bin (very) high, but it’s hard to tell by
eye if that is significant.</p>

<h2 id="significance">Significance</h2>

<p>To check whether the dice are biased, a <a href="https://en.wikipedia.org/wiki/Pearson%27s_chi-squared_test">chi-squared test</a> is required.
The chi-squared test essentially looks at how far away each point in a
distribution is from the expected value for that point, and normalizes by the
variance. The test statistic is then compared to the results expected from a
<a href="https://en.wikipedia.org/wiki/Chi-squared_distribution">chi-squared distribution</a> and a significance is obtained. Running
this test on our dice yields the following results:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Dice</th>
      <th style="text-align: right">chi-squared</th>
      <th style="text-align: right"><em>p</em>-value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Blue</td>
      <td style="text-align: right">26.31</td>
      <td style="text-align: right">0.001</td>
    </tr>
    <tr>
      <td style="text-align: left">Black</td>
      <td style="text-align: right">9.32</td>
      <td style="text-align: right">0.315</td>
    </tr>
    <tr>
      <td style="text-align: left">Red</td>
      <td style="text-align: right">10.77</td>
      <td style="text-align: right">0.215</td>
    </tr>
    <tr>
      <td style="text-align: left">White</td>
      <td style="text-align: right">19.07</td>
      <td style="text-align: right">0.014</td>
    </tr>
  </tbody>
</table>

<p>The chi-squared test has some <a href="https://stats.stackexchange.com/q/93212">caveats about low expected values</a>,
but at worst we only have two (out of nine) bins below five expected entries.
Looking at the <a href="https://en.wikipedia.org/wiki/p-value"><em>p</em>-values</a> we conclude roughly the same as our
“<em>chi-by-eye</em>” test above: the blue dice are significantly biased, while the
black and red dice show no evidence of being unfair. The white dice are not
biased at the <em>p</em> &lt; 0.01 level, but that single high bin is odd and to be
absolutely sure I would want to roll them a lot more and check.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="fun-and-games" />
        
      

      

      
      
        <summary type="html"><![CDATA[My friends and I played a Fate RPG for over two years. During that time we rolled a lot of dice and developed a lot of superstitions, but were any of them correct?]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/fate-dice-statistics/fate_dice.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/fate-dice-statistics/fate_dice.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Caltrain Visual Schedule</title>
      <link href="https://alexgude.com/blog/caltrain-visual-schedule/" rel="alternate" type="text/html" title="Caltrain Visual Schedule" />
      <published>2017-05-21T00:00:00-07:00</published>
      <updated>2017-05-21T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/caltrain_visual_schedule</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/caltrain-visual-schedule/"><![CDATA[<p>In 1878, <a href="https://en.wikipedia.org/wiki/%C3%89tienne-Jules_Marey">Étienne-Jules Marey</a> published <a href="https://archive.org/details/lamthodegraphiq00maregoog"><em>La Méthode
Graphique</em></a>,<sup style="anchor-name:--fnref-lmgd" id="fnref:lmgd"><a href="#fn:lmgd" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> a manual on using graphs for data
analysis. The book included Ibry’s<sup style="anchor-name:--fnref-ibry" id="fnref:ibry"><a href="#fn:ibry" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> <a href="/files/caltrain-schedule/ibry-trainschedule.jpg">famous visualization of a French
train schedule</a> which shows the position (y-axis) of trains
traveling from Paris to Lyon as a function of the time of day (x-axis). The
schedule elegantly packs a lot of information into a small space: the speed
and direction of trains are indicated by their slope, and when lines cross it
indicates that the trains pass each other. The schedule is such an iconic
visualization that <a href="https://en.wikipedia.org/wiki/Edward_Tufte">Tufte</a> used it as the cover of <em>The Visual Display
of Quantitative Information</em>.</p>

<p><a href="/files/caltrain-schedule/ibry-trainschedule.jpg"><img src="/files/caltrain-schedule/ibry-trainschedule.jpg" alt="Graph showing the progress of trains on a railway, according to the method
of Ibry" /></a></p>

<p>While Ibry’s schedule is the most famous example, it was not the first: <a href="https://dx.doi.org/10.1080/09332480.2013.772394">an
earlier example was produced by Serjev in Russia</a>. Nor was it the last,
as many people have produced similar diagrams for <a href="https://mbtaviz.github.io/">the T in Boston</a>,
<a href="https://web.archive.org/web/20220402191513/http://www.drones.com/bart.html">BART</a>, and <a href="https://web.archive.org/web/20180827045956/http://vis.berkeley.edu/courses/cs294-10-sp10/wiki/index.php/A4-PaulIvanov">Caltrain</a> (and <a href="https://mbostock.github.io/protovis/ex/caltrain-full.html">over</a>, and
<a href="https://www.davidstarke.com/projects/caltrain/">over</a>, and <a href="https://web.archive.org/web/20260716165216/https://www.svds.com/wp-content/uploads/2016/05/DataEDGE_2016.pdf#page=14">over</a> again). As a frequent
<a href="https://en.wikipedia.org/wiki/Caltrain">Caltrain</a> commuter, I thought I would try to put my spin on it. You
can find the Jupyter notebook used to make these schedules <a href="/files/caltrain-schedule/Caltrain%20Marey%20Schedule.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/caltrain-schedule/Caltrain%20Marey%20Schedule.ipynb">rendered on Github</a>). The schedule data is from <a href="http://www.caltrain.com/developer.html">Caltrain’s
developer site</a>.</p>

<h2 id="caltrain">Caltrain</h2>

<p>The three types of train are color coded as follows: local trains are blue,
limited-stop trains are green, and baby bullets are red. Every stop for a
train is indicated by a circle. The spacing between the stations on the y-axis
is scaled to the actual distance <a href="https://en.wikipedia.org/wiki/List_of_Caltrain_stations">recorded on the track
mileposts</a>.<sup style="anchor-name:--fnref-note" id="fnref:note"><a href="#fn:note" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> Click the schedules for larger versions.</p>

<h3 id="weekday">Weekday</h3>

<p>On the weekday there are so many trains that the full schedule is hard to
read, so instead I have focused on the morning and evening commute times. The
full weekday schedule is <a href="/files/caltrain-schedule/caltrain_weekday.svg">here</a> (<a href="/files/caltrain-schedule/caltrain_weekday_north.svg">northbound only</a> and
<a href="/files/caltrain-schedule/caltrain_weekday_south.svg">southbound only</a>). The full schedule with Gilroy included is
<a href="/files/caltrain-schedule/caltrain_weekday_gilroy.svg">here</a>.</p>

<p><a href="/files/caltrain-schedule/caltrain_weekday_morning.svg"><img src="/files/caltrain-schedule/caltrain_weekday_morning.svg" alt="Marey visual train schedule for caltrain during the weekday morning commute" /></a>
<a href="/files/caltrain-schedule/caltrain_weekday_evening.svg"><img src="/files/caltrain-schedule/caltrain_weekday_evening.svg" alt="Marey visual train schedule for caltrain during the weekday evening commute" /></a></p>

<p>In the morning there are six northbound bullets and five southbound while in
the evening the numbers are reversed. There are pairs of limited trains where
one of the pair makes most southern stops, and the other makes most northern
stops; the train making fewer stops initially catches up but then falls back
again as the lead train starts making fewer stops. We can see that the bullets
overtake the limited and local trains <a href="https://en.wikipedia.org/wiki/Caltrain_Express">near Bayshore and Lawrence</a>.
Finally, it is tough to see because I have cut off stations after
<a href="https://en.wikipedia.org/wiki/Tamien_Station">Tamien</a>, but three trains head north from <a href="https://en.wikipedia.org/wiki/Gilroy_station">Gilroy</a> early in
the morning, and three trains end there in the evening, ready for the next
morning’s commute.</p>

<h3 id="weekend">Weekend</h3>

<p><a href="/files/caltrain-schedule/caltrain_saturday.svg"><img src="/files/caltrain-schedule/caltrain_saturday.svg" alt="Marey visual train schedule for caltrain on Saturday" /></a>
<a href="/files/caltrain-schedule/caltrain_sunday.svg"><img src="/files/caltrain-schedule/caltrain_sunday.svg" alt="Marey visual train schedule for caltrain on Sunday" /></a></p>

<p>The weekends have far fewer trains. A local train runs in each direction
hourly, and there are four bullets each day. Only one bullet is running at a
time, and so it is possible they use the same <a href="https://en.wikipedia.org/wiki/Rolling_stock">rolling stock</a> for the
north and southbound trips. Saturday has two more north bound, and one more
southbound train than Sunday.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:lmgd">
      <p>The full title is <em>La Méthode Graphique Dans les Sciences Expérimentales et Principalement en Physiologie et en Médecine</em>, or roughly <em>The Graphical Method in Experimental Sciences and Mainly in Physiology and Medicine</em>. <a href="#fnref:lmgd" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ibry">
      <p>The caption in Marey’s book reads: <em>“Graphique de la marche des trains sur un chemin de fer, d’après la méthode de Ibry”</em> or <em>“Graph showing the progress of trains on a railway, according to the method of Ibry”</em>. Unfortunately, little else is known of Ibry, and so this type of chart is often named for Marey instead. <a href="#fnref:ibry" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:note">
      <p>The mileposts markers are off by up to 100m for stations south of <a href="https://en.wikipedia.org/wiki/Lawrence_station_(Caltrain)">Lawrence</a>. I tried to measure the actual track distances using <a href="https://www.openstreetmap.org/#map=11/37.5574/-122.3050&amp;layers=T">OpenStreetMap</a> but found that I could not do so more accurately than the milepost numbers. <a href="#fnref:note" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[In 1878, Marey published a famous visual train schedule based on work by Ibry. What would it look like for Silicon Valley's Caltrain? Come find out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/caltrain-schedule/sp_3208_with_train_128_in_redwood_city_ca_in_august_1980_by_roger_puta.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/caltrain-schedule/sp_3208_with_train_128_in_redwood_city_ca_in_august_1980_by_roger_puta.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Visualizing Multiple Data Distributions</title>
      <link href="https://alexgude.com/blog/distribution-plots/" rel="alternate" type="text/html" title="Visualizing Multiple Data Distributions" />
      <published>2017-04-24T00:00:00-07:00</published>
      <updated>2017-04-24T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/distribution_plots</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/distribution-plots/"><![CDATA[<p>One of the first steps when exploring data is to look at its distribution. For
single distributions, or for comparing a small number, <a href="https://en.wikipedia.org/wiki/Histogram">histograms</a> are
great. However, as the number of distributions to compare grows, histograms
become less and less useful for visualizing data. Fortunately, there are some
good alternatives.</p>

<p>In this post, I’ll look at a few different plot types I explored when
comparing the <a href="/blog/switrs-crashes-by-date/#day-of-the-week">distributions of crashes by day of the week</a>. More
information about the data can be found in the original post: <a href="/blog/switrs-crashes-by-date/"><em>SWITRS: On
What Days Do People Crash?</em></a></p>

<p>The Jupyter notebook used to make these plots can be found <a href="/files/distribution-plots/Plot%20Types.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/distribution-plots/Plot%20Types.ipynb">rendered on Github</a>).</p>

<h2 id="box-plots">Box Plots</h2>

<p><a href="https://en.wikipedia.org/wiki/Box_plot"><strong>Box Plots</strong></a>, or box-and-whisker plots, are one of the simpler ways of
plotting a series of distributions. The edges of the box show the 1st and 3rd quartile
while the line within the box shows the median (2nd quartile). The whiskers
show the extent of the data, but their usage is not standardized. Sometimes
they show the full extent of the data, sometimes some percentage of the
inner-quartile range, and sometimes one standard deviation from the mean. If
they do not show the full extent, the points not included in the whiskers are
plotted individually.</p>

<p>Box plots quickly convey some essential statistics about the distributions
and make no assumptions about the underlying data, which is both a strength
and a weakness. Their simplicity can hide important information, and their
non-standard whiskers can cause confusion if what they represent is not
clearly stated.</p>

<p><a href="/files/distribution-plots/accidents_by_day_of_the_week_box.svg"><img src="/files/distribution-plots/accidents_by_day_of_the_week_box.svg" alt="A box plot showing the distribution of crashes per day in California from
2001–2015 by day of the week." /></a></p>

<p>These box plots show the distributions of the <a href="/blog/switrs-crashes-by-date/#day-of-the-week">number of crashes per day in
California by day of the week</a> from 2001–2015. From the box plots it is
easy to see that there are more crashes on Fridays and that the weekends have
fewer crashes than the weekdays. Most of the outliers are on the high side,
but we can’t tell anything about the actual shape of the distributions.</p>

<h2 id="strip-plots">Strip Plots</h2>

<p><a href="https://en.wikipedia.org/wiki/Dot_plot_(statistics)#Dot_plots"><strong>Strip Plots</strong></a>, also called dot plots or univariate dot plots, try
to give us a little—well a lot—more information than box plots. They plot
<em>every</em> point in the dataset, which can give you a good view of what is
happening. They often have a bit of random jitter added to each point along
the categorical axis so that the points do not overlap as much.</p>

<p>Strip plots make it is easy to see all the outliers. The density of points
also gives an approximation of the underlying distribution, although this can
be hard to judge by eye because the distance in the categorical axis, while
meaningless, obscures the true distance between points. Overlapping points
also make it tough to estimate the true distribution, especially as the number
of points increases.</p>

<p><a href="/files/distribution-plots/accidents_by_day_of_the_week_strip.svg"><img src="/files/distribution-plots/accidents_by_day_of_the_week_strip.png" alt="A strip plot showing the distribution of crashes per day in California
from 2001–2015 by day of the week." /></a></p>

<p>With the strip plot it is still possible to tell which days have more
crashes, but without the quartiles to guide the eye it is not as easy. We
can now start to see that the distributions are bimodal, although it is
difficult to see details with all the points. Strip plots are enticing because
they show literally all of the data, but plots which summarize the data are
often more useful for large datasets.</p>

<h2 id="swarm-plots">Swarm Plots</h2>

<p><a href="https://cran.r-project.org/package=beeswarm"><strong>Swarm Plots</strong></a>, also called beeswarm plots, are similar to strip
plots in that they plot all of the data points. Unlike strip plots, swarm
plots attempt to avoid obscuring points by calculating non-overlapping
positions instead of adding random jitter. This sort of gives them the appearance
of a swarm of bees, or perhaps a honeycomb.</p>

<p>Swarm plots share many of the same advantages of strip plots, but without as
much clutter to hide their salient features. Unfortunately, spreading out the
points in a non-overlapping fashion limits the number of points that can be
plotted—there is only so much space on the page! Additionally, the algorithm
that calculates the positions is computationally expensive and so scales poorly
as the number of points increases.<sup style="anchor-name:--fnref-time" id="fnref:time"><a href="#fn:time" class="footnote" rel="footnote" role="doc-noteref">1</a></sup></p>

<p>This slow generation time is especially harmful during exploratory analyses.
It is easy to keep engaged with the problem when a plot takes a second
or two to pop up, but when they take more than a minute my productivity
plummets as my iteration time explodes and I have to constantly context
switch.</p>

<p><a href="/files/distribution-plots/accidents_by_day_of_the_week_swarm.svg"><img src="/files/distribution-plots/accidents_by_day_of_the_week_swarm.png" alt="A swarm plot showing the distribution of crashes per day in California
from 2001–2015 by day of the week." /></a></p>

<p>I generated this plot using a sampled subset of the data because the swarms
piled up when trying to show the full dataset. Even so, you can see some of the
points have piled up against the edges of each column. The swarm plot makes
the relative crash rates easier to see than on the strip plot. The bimodal
nature of the distributions is much clearer and their shape can almost be made
out. However, the thickness of the plotted points causes the formation of the
strands extending out from each swarm which make judging the true shape of the
distributions difficult.</p>

<h2 id="violin-plots">Violin Plots</h2>

<p><a href="https://en.wikipedia.org/wiki/Violin_plot"><strong>Violin plots</strong></a> try to give an indication of the distribution
of data without cluttering the plot by drawing all of the points. They do
this by using <a href="https://en.wikipedia.org/wiki/Kernel_density_estimation">kernel density estimation (KDE)</a> to model the
distribution. There is a lot of information to process in a violin plot so
they can be a bit tough to read. The shape of the violin body indicates the
number of observations: if the violin is thick at some value it means there
are a lot of data points there, if it is thin then there are few. The inside
of the violin is often marked to indicate additional information. The violins
below have the quartiles drawn inside them as dashed lines, but miniature box
plots are another common inner marking.</p>

<p>The main disadvantage of violin plots is that the KDE bandwidth must be
selected. Too low and the features of the data are washed out. Too high and
the KDE overfits the data. This limits their usefulness when there are only a
few data points. The lack of standardization when it comes to the inner
markings also makes them hard to interpret if they aren’t explicitly
explained.</p>

<p><a href="/files/distribution-plots/accidents_by_day_of_the_week_violin.svg"><img src="/files/distribution-plots/accidents_by_day_of_the_week_violin.svg" alt="A violin plot showing the distribution of crashes per day in California
from 2001–2015 by day of the week." /></a></p>

<p>The violin plots make the bimodal nature of the distributions crystal clear.
Likewise it is easy to see that there is an increase on Friday and a decrease
on Sunday. However, we have lost sight of our outliers. We can see that the
Friday violin extends to almost 3000 crashes, but exactly how many data
points go into that thin tail is unclear.</p>

<h2 id="conclusion">Conclusion</h2>

<p>I like violin plots <a href="/blog/switrs-crashes-by-date/#day-of-the-week"><strong>a</strong></a> <a href="/blog/switrs-motorcycle-crashes-by-date/#day-of-the-week"><strong>whole</strong></a> <a href="/blog/switrs-daylight-saving-time-accidents/#crash-ratio"><strong>lot</strong></a>! While
swarm and strip plots show lots of detail and box plots provide good overviews
with the summary statistics, I find violin plots to be a good middle ground.
The KDE provides more detail than a pure box plot, includes the same useful
summary statistics, and avoids cluttering the plot with every data point.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:time">
      <p>The box and violin plots in this post take about a second to render on my desktop. The strip plot take 5 seconds. The swarm plot take <strong>74 seconds</strong>! <a href="#fnref:time" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="data-visualization" />
        
      

      

      
      
        <summary type="html"><![CDATA[Need to compare a set of distributions of some variable? Histograms are OK, but try something fancier! Read on to learn about box, strip, swarm, and violin plots!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/distribution-plots/Petrov-Vodkin_violin_1921.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/distribution-plots/Petrov-Vodkin_violin_1921.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: Car Crashes After Daylight Saving Time</title>
      <link href="https://alexgude.com/blog/switrs-daylight-saving-time-accidents/" rel="alternate" type="text/html" title="SWITRS: Car Crashes After Daylight Saving Time" />
      <published>2017-03-20T00:00:00-07:00</published>
      <updated>2017-03-20T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/switrs_daylight_saving_time_accidents</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-daylight-saving-time-accidents/"><![CDATA[<p>The <a href="https://en.wikipedia.org/wiki/Daylight_saving_time">daylight saving time</a> (DST) change is awful—we get less sleep and
it <a href="https://www.scientificamerican.com/article/does-daylight-saving-times-save-energy/">might not even save energy</a> as was intended! Worse, studies by
<a href="https://doi.org/10.1016/S1389-9457(00)00032-0">Varughese &amp; Allen</a><sup style="anchor-name:--fnref-varughese_cite" id="fnref:varughese_cite"><a href="#fn:varughese_cite" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and
<a href="https://doi.org/10.1257/app.20140100">Smith</a><sup style="anchor-name:--fnref-smith_cite" id="fnref:smith_cite"><a href="#fn:smith_cite" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> have shown that the time change increases the
number of automobile crashes! Let’s look for a similar trend in the <a href="/blog/switrs-to-sqlite/">SWITRS
data</a> that I’ve collected.</p>

<p>The Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-dst/SWITRS%20Daylight%20Saving%20Time%20Crashes.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-dst/SWITRS%20Daylight%20Saving%20Time%20Crashes.ipynb">rendered on Github</a>).</p>

<h2 id="crash-ratio">Crash Ratio</h2>

<p>The analysis is relatively simple. I start with the number of crashes that
happen on the days following the start of DST in California. I divide the
amount of crashes on each day by the number of crashes on the same day of the
week but two weeks later.<sup style="anchor-name:--fnref-after" id="fnref:after"><a href="#fn:after" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> Taking the ratio cancels out most of the effects
that are unrelated to the time change—like the fact that <a href="/blog/switrs-crashes-by-date/#crashes-per-week">crash rates vary
by 30% depending on the year</a>. Two weeks after is a good choice for
normalization because:</p>

<ul>
  <li>
    <p>The weeks after the time change have similar daylight hours to the week of
the time change.</p>
  </li>
  <li>
    <p>The crashes rate is still slightly elevated a week later, so normalizing by
the very next week hides some of the increase that is due to the start of
DST.<sup style="anchor-name:--fnref-back_to_normal" id="fnref:back_to_normal"><a href="#fn:back_to_normal" class="footnote" rel="footnote" role="doc-noteref">4</a></sup></p>
  </li>
</ul>

<p>The <a href="https://en.wikipedia.org/wiki/Violin_plot">violin plots</a> below show the distribution of these ratios from
the years 2001 to 2016. A value greater than 1 means that there are more
crashes during the week when DST starts than two weeks after.</p>

<p><a href="/files/switrs-dst/accidents_two_weeks_after_dst_change_in_california.svg"><img src="/files/switrs-dst/accidents_two_weeks_after_dst_change_in_california.svg" alt="Violin plot showing the ratio of crashes per day of the week for the week
after the start of daylight saving time, divided by the week two weeks
after." /></a></p>

<p>Except for Sunday, every day of the week following the time change has on
average a higher rate of crashes! I am surprised that the crash rate stays
high the entire week. This indicates that it takes even longer than a week for
people to catch up on sleep and for the crash rate to go back to normal.</p>

<h2 id="t-test"><em>t</em>-Test</h2>

<p>So the “<em>chi-by-eye</em>” plot is suggestive, but I can quantify whether the
results are significant using a <a href="https://en.wikipedia.org/wiki/Student%27s_t-test#Paired_samples">two-tailed paired <em>t</em>-test</a>.
This is the method Varughese &amp; Allen use. Doing so gives the following
results.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Day</th>
      <th style="text-align: right"><em>t</em>-value</th>
      <th style="text-align: right"><em>p</em>-value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><strong>Monday</strong></td>
      <td style="text-align: right">2.7</td>
      <td style="text-align: right"><strong>0.017</strong></td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Tuesday</strong></td>
      <td style="text-align: right">2.5</td>
      <td style="text-align: right"><strong>0.023</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Wednesday</td>
      <td style="text-align: right">1.6</td>
      <td style="text-align: right">0.122</td>
    </tr>
    <tr>
      <td style="text-align: left">Thursday</td>
      <td style="text-align: right">1.4</td>
      <td style="text-align: right">0.191</td>
    </tr>
    <tr>
      <td style="text-align: left">Friday</td>
      <td style="text-align: right">0.9</td>
      <td style="text-align: right">0.361</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>Saturday</strong></td>
      <td style="text-align: right">2.6</td>
      <td style="text-align: right"><strong>0.019</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">Sunday (DST)</td>
      <td style="text-align: right">-1.6</td>
      <td style="text-align: right">0.121</td>
    </tr>
  </tbody>
</table>

<p>The increase in crashes on Monday is significant, as is Tuesday and Saturday.
Sunday is the only day that trends lower (matching the plot), but not
significantly.</p>

<p>So daylight savings time causes more crashes, but those of us in California
might be in luck! State Assembly member <a href="https://en.wikipedia.org/wiki/Kansen_Chu">Kansen Chu</a> has introduced a
bill to <a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201520160AB385">finally do away with DST</a>! Hopefully it will pass and let us
all get that hour of sleep we deserve.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:varughese_cite">

      <p><span class="citation">Varughese, J. and Allen, R. <a href="https://doi.org/10.1016/S1389-9457(00)00032-0">“Fatal accidents following changes in daylight savings time: the American experience”</a> <cite>Sleep Medicine</cite>. vol. 2, no. 1. 2000. pp. 31–36. doi: <a href="https://doi.org/10.1016/S1389-9457(00)00032-0">10.1016/S1389-9457(00)00032-0</a>.</span> <a href="#fnref:varughese_cite" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:smith_cite">

      <p><span class="citation">Smith, Austin C. <a href="https://doi.org/10.1257/app.20140100">“Spring Forward at Your Own Risk: Daylight Saving Time and Fatal Vehicle Crashes”</a> <cite>American Economic Journal: Applied Economics</cite>. vol. 8, no. 2. 2016. pp. 65–91. doi: <a href="https://doi.org/10.1257/app.20140100">10.1257/app.20140100</a>.</span> <a href="#fnref:smith_cite" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:after">

      <p>It is also possible to use the week before or the week directly
after the DST change to normalize. For the curious, I have also
made <a href="/files/switrs-dst/accidents_after_dst_change_in_california_before.svg">a plot using the week before for normalization</a>
and <a href="/files/switrs-dst/accidents_after_dst_change_in_california.svg">the week after</a>. They both show the same trend. <a href="#fnref:after" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:back_to_normal">

      <p>I assume that people are back to normal after three weeks, and so
I use that week as a control. I then compare that ratios of the
control week with <a href="/files/switrs-dst/accidents_one_and_three_weeks_after_dst_change_in_california.svg">one week after the DST change</a> and
<a href="/files/switrs-dst/accidents_two_and_three_weeks_after_dst_change_in_california.svg">two weeks after the DST change</a> to see which is more
normal. One week after has Monday and Thursday high, indicating
people are still having more crashes than we expect. Two weeks
after the ratios are near one, and so I conclude people are back
to normal by then. <a href="#fnref:back_to_normal" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Daylight saving time leaves us drowsy and cranky at work, but it also leads to an increase in traffic collisions! Find out exactly how many more there are with this analysis!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-dst/dst_change_gare_saint_lazare_1937.png" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-dst/dst_change_gare_saint_lazare_1937.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: On What Days Do Motorcycles Crash?</title>
      <link href="https://alexgude.com/blog/switrs-motorcycle-crashes-by-date/" rel="alternate" type="text/html" title="SWITRS: On What Days Do Motorcycles Crash?" />
      <published>2017-02-21T00:00:00-08:00</published>
      <updated>2017-02-21T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_motorcycle_crashes_by_date</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-motorcycle-crashes-by-date/"><![CDATA[<p>A few months ago I wrote a post in which <a href="/blog/switrs-crashes-by-date/">I explored when car crashes happen
in California</a>. This time I’m going to go through the same analysis
but restrict myself to looking at crashes involving motorcycles. Motorcycle
crashes are the original reason I tracked down the <a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">SWITRS</a> data: my
father rode motorcycles for years (he only recently stopped) and we wanted to
better understand what sort of risks that brought.</p>

<p>I expected the crash trend for motorcycles to match the one I found when
<a href="/blog/switrs-crashes-by-date/">looking at cars</a>. There I found that commute crashes accounted for
the majority of crashes, and so holidays and weekends that most people have
off result in fewer crashes. Motorcycles, we will see, do not follow this
pattern.</p>

<p>One thing before we get started: the number of riders on a given day (or more
accurately, the <a href="https://en.wikipedia.org/wiki/Vehicle_miles_of_travel">number of miles ridden by them</a>) has the most impact on
the number of crashes. If there are more riders, there are going to be more
crashes. From now on I’ll treat the two numbers as equivalent, even though
there are some confounding factors, like weather, which would change the ratio
of crashes to number of riders on the road; I hope to look at these other
factors in a later post.</p>

<p>As per usual, the Jupyter notebook used to perform this analysis can be found
<a href="/files/switrs-motorcycle-accidents-by-date/SWITRS%20Crash%20Dates%20With%20Motorcycles.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-motorcycle-accidents-by-date/SWITRS%20Crash%20Dates%20With%20Motorcycles.ipynb">rendered on Github</a>).</p>

<h2 id="data-selection">Data Selection</h2>

<p>I selected crashes involving motorcycles from the <a href="https://github.com/agude/SWITRS-to-SQLite">SQLite database</a>
(<a href="/blog/switrs-to-sqlite/">discussed previously</a>) with the following query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">Collision_Date</span> <span class="k">FROM</span> <span class="n">Collision</span>
<span class="k">WHERE</span> <span class="n">Collision_Date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">Motorcycle_Collision</span> <span class="o">==</span> <span class="mi">1</span>       <span class="c1">-- Involves a motorcycle</span>
<span class="k">AND</span> <span class="n">Collision_Date</span> <span class="o">&lt;=</span> <span class="s1">'2015-12-31'</span>  <span class="c1">-- 2016 is incomplete</span>
</code></pre></div></div>

<p>This gave me 193,336 data points (crashes) to examine spanning 2001 through
2015. <a href="/blog/switrs-crashes-by-date/#data-selection">Just as before</a>, crashes from 2016 are rejected because there is
not yet complete data for the year.</p>

<h2 id="crashes-per-week">Crashes per Week</h2>

<p>For cars, <a href="/blog/switrs-crashes-by-date/#crashes-per-week">I found that there was a decrease in crashes</a> starting in 2008
as people stopped driving to work during the <a href="https://en.wikipedia.org/wiki/Great_Recession">Great Recession</a>. Apart from
that, I found that the week-to-week rate changed relatively little, with
holidays providing the largest increases and decreases. When I looked at just
motorcycles, I expected to see a similar pattern. However, the trends (plotted
below) are completely different.</p>

<p><a href="/files/switrs-motorcycle-accidents-by-date/motorcycle_accidents_per_week_in_california.svg"><img src="/files/switrs-motorcycle-accidents-by-date/motorcycle_accidents_per_week_in_california.svg" alt="Line plot showing crashes per week from 2001 to
2015" /></a></p>

<p>There are far fewer crashes because there are far fewer motorcycles; there are
about <a href="https://www.fhwa.dot.gov/policyinformation/statistics/2012/mv1.cfm">27 million vehicles in California, but of those only 770,000 are
motorcycles</a>. There is also a strong seasonal effect—even in sunny
California, motorcycle ridership drops drastically in the winter! And unlike
for cars, there is not a large decrease due to the recession. Finally, there
is an overall upward trend in the number of motorcycle crashes.</p>

<p>Commute crashes account for the majority of car crashes. However, this does
not appear to be the case for motorcycles, because the numbers were relatively
unchanged by the Great Recession; people kept riding at the same levels when
out of work.</p>

<h2 id="day-by-day">Day-by-Day</h2>

<p>For car, <a href="/blog/switrs-crashes-by-date/#day-by-day">the days with the most and least crashes were holidays</a>. The
largest number of crashes were on holidays where people went to work and then
out afterwards, like Halloween. Motorcycle crashes do not follow that pattern.
Instead, the holidays show quite disparate results: some holidays dip, some
spike, others show almost no deviation from a normal day.</p>

<p><a href="/files/switrs-motorcycle-accidents-by-date/mean_motorcycle_accidents_by_date.svg"><img src="/files/switrs-motorcycle-accidents-by-date/mean_motorcycle_accidents_by_date.svg" alt="Line plot showing average motorcycle crashes by day of the
year" /></a></p>

<p>The summer holidays do not stand out; only Memorial Day is readily visible.
Winter holidays, by contrast, show both peaks and valleys. I would interpret
this as due to the seasonal weather:</p>

<ul>
  <li>
    <p>In summer, any day is a good day to ride.</p>
  </li>
  <li>
    <p>In the winter, the weather keeps riders off the road, except when a holiday
gives them the extra motivation they need.</p>
  </li>
</ul>

<p>One final outlier to address: the sharp peak at the end of February is <a href="https://en.wikipedia.org/wiki/February_29">leap
day</a>. The peak is not an error, but is a statistical artifact. The mean
for all other days is calculated with <code class="language-plaintext highlighter-rouge">n = 15</code>, but only <code class="language-plaintext highlighter-rouge">n = 3</code> for leap day.</p>

<h2 id="day-of-the-week">Day of the Week</h2>

<p>Car crashes <a href="/blog/switrs-crashes-by-date/#day-of-the-week">happen less during the weekend, when people aren’t
commuting</a>. For motorcycles the weekends are prime riding times, and so
the number of crashes increases.</p>

<p>If we think of weekends as a kind of mini-holiday, they provide a way to look
at the same seasonal holiday phenomenon <a href="#day-by-day">discussed above</a>. Winter
holidays showed high variance, so I would expect to see some weekends with
high winter ridership, and some with low ridership. Summer holidays had low
variance, so I expect to see similar ridership on all summer weekends.</p>

<p>The <a href="https://en.wikipedia.org/wiki/Violin_plot">violin plots</a> below show the distribution of crashes by day of
the week over the 15 year period. They are divided into two seasons: summer
(May–October) and winter (November–April).</p>

<p><a href="/files/switrs-motorcycle-accidents-by-date/motorcycle_accidents_by_day_of_the_week_and_season.svg"><img src="/files/switrs-motorcycle-accidents-by-date/motorcycle_accidents_by_day_of_the_week_and_season.svg" alt="Violin plot showing crashes by day of the week in summer and
winter" /></a></p>

<p>There is lower ridership in winter over all (top row), as indicated by the
central dotted line indicating average number of crashes. And we can see an
increase on weekends; but during the winter, that weekend increase is small as
compared with summer (bottom row). The winter distributions are more elongated
than those from summer, meaning that on some days there are many riders, and
on others there are almost none, just as we expected. Summer weekends, by
contrast, have consistently high ridership.</p>

<p>We can conclude that weekend rider behavior does seem to track seasonal
holiday riding behavior. And like the trends for holidays, the weekend results
could be due to weather.</p>

<h2 id="conclusion">Conclusion</h2>

<p>Motorcycle crashes do not follow the same trends as for cars. Motorcyclists
continue riding even when they do not have a job to commute to. Seasons have a
large effect on the number of riders out on the road. Motorcycle ridership has
variance for winter holidays and weekends when the weather may turn against
them. There are many more ways to explore motorcycle crashes—time of day,
type of motorcycle, vehicle at fault—but those will have to wait for another
day.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Motorcycles riders are a different breed, born to chase excitement! So when do they crash? Using California's SWITRS data I find out! I'll give you a hint: it is not on the way to their 9-5!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-motorcycle-accidents-by-date/police_in_stockholm.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-motorcycle-accidents-by-date/police_in_stockholm.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Software Testing for Data Science</title>
      <link href="https://alexgude.com/blog/software-testing-for-data-science/" rel="alternate" type="text/html" title="Software Testing for Data Science" />
      <published>2017-01-12T00:00:00-08:00</published>
      <updated>2017-01-12T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/software_testing_for_data_science</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/software-testing-for-data-science/"><![CDATA[<p>From the title <em>“Data Scientist”</em> you would guess that we work with data, and
we do some sort of science, but what you might miss is that we write a lot of
code. This means that good development practices <strong>are</strong> good data
science practices! A good practice to start with, one that is not only easy to
do but extremely useful, is <a href="https://en.wikipedia.org/wiki/Software_testing"><strong>testing</strong></a>.</p>

<p>Working tests provide several benefits for your data science projects:</p>

<ul>
  <li>
    <p>They give you confidence that your code works as intended.</p>
  </li>
  <li>
    <p>They allow rapid development because breaks introduced by new code are
caught sooner.</p>
  </li>
  <li>
    <p>They encode the assumptions you have about how your code works in a standard
form.</p>
  </li>
</ul>

<p>Below I’ll go through a few ways to test data science code using examples from
my <a href="https://github.com/agude/SWITRS-to-SQLite">SWITRS-to-SQLite</a> project (which I’ve <a href="/blog/switrs-to-sqlite/">discussed
previously</a>).</p>

<h2 id="data-extraction-tests">Data Extraction Tests</h2>

<p>The first step in a data driven analysis is reading your data. Often data is
read from a database, but sometimes (as was the case in SWITRS) it is from a
file on disk. Testing your data loading involves generating a fake data source
and attempting to read it. For example, <a href="https://github.com/agude/SWITRS-to-SQLite/blob/5167c7d9ffe7384224a76b1f209b7638fdd70362/tests/test_open_records_file.py#L8-L18">here</a> is the test which
verifies that my code can read <code class="language-plaintext highlighter-rouge">.gzip</code> data files:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">etl</span> <span class="kn">import</span> <span class="n">open_record_file</span>

<span class="k">def</span> <span class="nf">test_read_gzipped_file</span><span class="p">(</span><span class="n">tmpdir</span><span class="p">):</span>
    <span class="c1"># Write a file to read back
</span>    <span class="n">f</span> <span class="o">=</span> <span class="n">tmpdir</span><span class="p">.</span><span class="nf">join</span><span class="p">(</span><span class="sh">"</span><span class="s">test.csv.gz</span><span class="sh">"</span><span class="p">)</span>
    <span class="n">contents</span> <span class="o">=</span> <span class="sh">"</span><span class="s">Test contents</span><span class="se">\n</span><span class="s">second line</span><span class="sh">"</span>
    <span class="n">file_path</span> <span class="o">=</span> <span class="n">f</span><span class="p">.</span><span class="n">dirname</span> <span class="o">+</span> <span class="sh">"</span><span class="s">/</span><span class="sh">"</span> <span class="o">+</span> <span class="n">f</span><span class="p">.</span><span class="n">basename</span>
    <span class="k">with</span> <span class="n">gzip</span><span class="p">.</span><span class="nf">open</span><span class="p">(</span><span class="n">file_path</span><span class="p">,</span> <span class="sh">'</span><span class="s">wt</span><span class="sh">'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="n">f</span><span class="p">.</span><span class="nf">write</span><span class="p">(</span><span class="n">contents</span><span class="p">)</span>

    <span class="c1"># Read back the file
</span>    <span class="k">with</span> <span class="nf">open_record_file</span><span class="p">(</span><span class="n">file_path</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="k">assert</span> <span class="n">f</span><span class="p">.</span><span class="nf">read</span><span class="p">()</span> <span class="o">==</span> <span class="n">contents</span>
</code></pre></div></div>

<p>In this test I create a very small gzip file, read it with the function being
tested, and verify that it reports the same content that I wrote to the file.
It was helpful to have this function when writing the library because I wanted
to be able to read both zipped files and plain text (since compressing the
files can save several gigabytes of disk space). This test ensured I didn’t
break the ability to read one of these file types while working on the other.</p>

<h2 id="data-cleaning-and-transformation-tests">Data Cleaning and Transformation Tests</h2>

<p>After the raw data is loaded, it is often necessary to parse (read, clean, and
transform) it. I’ve found that the parser code benefits enormously from tests
for three reasons:</p>

<ul>
  <li>
    <p>It has to deal with high number of edge cases because data is never as clean
as we would like it to be.</p>
  </li>
  <li>
    <p>It often has the most bespoke code because everyone’s data is different.</p>
  </li>
  <li>
    <p>It must change when the raw data format changes, which is more likely than
not to be outside of your control.</p>
  </li>
</ul>

<p>Testing the parser is simple: create pairs of examples of raw data and
correctly parsed results, then test that the code produces the right result
when provided the raw data. You should have (at least) one of these pairs for
each variation of data you expect to encounter. For example, here is a
simplified version of the <a href="https://github.com/agude/SWITRS-to-SQLite/blob/5167c7d9ffe7384224a76b1f209b7638fdd70362/tests/test_converters.py#L7-L40">tests for my <code class="language-plaintext highlighter-rouge">convert</code> function</a>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">etl</span> <span class="kn">import</span> <span class="n">convert</span>

<span class="k">def</span> <span class="nf">test_convert</span><span class="p">():</span>
    <span class="n">convert_vals</span> <span class="o">=</span> <span class="p">(</span>
        <span class="c1"># Pass through
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">),</span>
        <span class="c1"># Standard dtypes
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="mi">9</span><span class="p">),</span>
        <span class="c1"># With spaces
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">9 </span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="mi">9</span><span class="p">),</span>
        <span class="p">(</span><span class="sh">"</span><span class="s"> 9</span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="mi">9</span><span class="p">),</span>
        <span class="c1"># Nulls that do nothing
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="p">[</span><span class="sh">""</span><span class="p">],</span> <span class="mi">9</span><span class="p">),</span>
        <span class="c1"># Nulls that return None
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="p">[</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">],</span> <span class="bp">None</span><span class="p">),</span>
        <span class="p">(</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="p">[</span><span class="sh">"</span><span class="s">9</span><span class="sh">"</span><span class="p">],</span> <span class="bp">None</span><span class="p">),</span>
        <span class="c1"># Conversion failure
</span>        <span class="p">(</span><span class="sh">"</span><span class="s">a</span><span class="sh">"</span><span class="p">,</span> <span class="nb">int</span><span class="p">,</span> <span class="bp">None</span><span class="p">,</span> <span class="bp">None</span><span class="p">),</span>
    <span class="p">)</span>
    <span class="k">for</span> <span class="n">val</span><span class="p">,</span> <span class="n">dtype</span><span class="p">,</span> <span class="n">nulls</span><span class="p">,</span> <span class="n">answer</span> <span class="ow">in</span> <span class="n">convert_vals</span><span class="p">:</span>
        <span class="k">assert</span> <span class="nf">convert</span><span class="p">(</span><span class="n">val</span><span class="o">=</span><span class="n">val</span><span class="p">,</span> <span class="n">dtype</span><span class="o">=</span><span class="n">dtype</span><span class="p">,</span> <span class="n">nulls</span><span class="o">=</span><span class="n">nulls</span><span class="p">)</span> <span class="o">==</span> <span class="n">answer</span>
</code></pre></div></div>

<p>The <a href="https://github.com/agude/SWITRS-to-SQLite/blob/5167c7d9ffe7384224a76b1f209b7638fdd70362/switrs_to_sqlite/switrs_to_sqlite.py#L25-L64"><code class="language-plaintext highlighter-rouge">convert</code> function</a> takes a string, a type, and a list of
values that should map to the <a href="https://en.wikipedia.org/wiki/Null_(SQL)"><code class="language-plaintext highlighter-rouge">NULL</code> value</a>. If the string is in
the <code class="language-plaintext highlighter-rouge">NULL</code> list then <code class="language-plaintext highlighter-rouge">None</code> should be returned, otherwise the string should be
converted to the type and returned. The tests provide example inputs and check
that they match the example output.</p>

<p>The SWITRS data files contained many edge cases—as is common in raw data!
These edge cases include: multiple input values that all mean <code class="language-plaintext highlighter-rouge">NULL</code>, multiple
input values that all map to <code class="language-plaintext highlighter-rouge">True</code> or <code class="language-plaintext highlighter-rouge">False</code>, a mix of strings and numbers
in the same column, and occasionally spaces prepended or appended to values. A
solid set of unit tests gave me confidence that the parsing functions worked
correctly and allowed me to track all of the edge cases as I discovered them.</p>

<h2 id="integration-tests">Integration Tests</h2>

<p>Once all components have been tested in isolation, you should test
the system as a whole. This ensures that the components, which we are already
confident in individually, are put together correctly into a working program.
For example, I <a href="https://github.com/agude/SWITRS-to-SQLite/blob/5167c7d9ffe7384224a76b1f209b7638fdd70362/tests/test_victimrow.py">test the class</a> that parses a row from one of the CSV
files using the <code class="language-plaintext highlighter-rouge">convert</code> function (and a few others) tested above as follows:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">etl</span> <span class="kn">import</span> <span class="n">VictimRow</span>

<span class="c1"># You should do this with a fixture, but this is a simple example
</span><span class="n">ROWS</span> <span class="o">=</span> <span class="p">(</span>
  <span class="p">(</span>
    <span class="c1">#Case_ID    Party_Number  Victim_Role  Victim_Sex  Victim_Age
</span>    <span class="c1"># Input
</span>    <span class="p">[</span><span class="sh">'</span><span class="s">097293</span><span class="sh">'</span><span class="p">,</span>  <span class="sh">'</span><span class="s">1</span><span class="sh">'</span><span class="p">,</span>          <span class="sh">'</span><span class="s">2</span><span class="sh">'</span><span class="p">,</span>         <span class="sh">'</span><span class="s">-</span><span class="sh">'</span><span class="p">,</span>        <span class="sh">'</span><span class="s">20</span><span class="sh">'</span><span class="p">,],</span>
    <span class="c1"># Output
</span>    <span class="p">[</span><span class="sh">'</span><span class="s">097293</span><span class="sh">'</span><span class="p">,</span>  <span class="mi">1</span><span class="p">,</span>            <span class="sh">'</span><span class="s">2</span><span class="sh">'</span><span class="p">,</span>         <span class="bp">None</span><span class="p">,</span>        <span class="mi">20</span><span class="p">,],</span>
  <span class="p">),</span>
<span class="p">)</span>

<span class="k">def</span> <span class="nf">test_victimrows</span><span class="p">():</span>
    <span class="k">for</span> <span class="n">row</span><span class="p">,</span> <span class="n">answer</span> <span class="ow">in</span> <span class="n">ROWS</span><span class="p">:</span>
        <span class="n">c</span> <span class="o">=</span> <span class="nc">VictimRow</span><span class="p">(</span><span class="n">row</span><span class="p">)</span>
        <span class="k">assert</span> <span class="n">c</span><span class="p">.</span><span class="n">values</span> <span class="o">==</span> <span class="n">answer</span>
</code></pre></div></div>

<p>I provide example rows from the CSV as input along with the expected correct
output. You’ll notice it has to read strings and convert them into strings,
integers, and <code class="language-plaintext highlighter-rouge">NULL</code>. Verifying this manually by hand would be quite tedious!</p>

<h2 id="tests-remembering-so-you-dont-have-to">Tests: Remembering So You Don’t Have To</h2>

<p>A final remark: tests are great because they keep track of the complexity
of your input data! When you find a new value in your data that your code
doesn’t handle correctly, add it to the tests! This does two things for you:</p>

<ul>
  <li>
    <p>You’ll know your code is fixed when the test starts passing.</p>
  </li>
  <li>
    <p>You will never mishandle that value again, regardless of what happens in the
future.</p>
  </li>
</ul>

<p>This takes a lot of mental load off you, allowing you to focus more on the
science, and less on the data parsing!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="data-science" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Much of data science involves writing code; for data cleaning, parsing, and modeling. Software tests can ensure that your code does what you think it does!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/data-science-testing/brick_header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/data-science-testing/brick_header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Eldar: A Bright, High-contrast Color Scheme for Vim</title>
      <link href="https://alexgude.com/blog/vim-eldar/" rel="alternate" type="text/html" title="Eldar: A Bright, High-contrast Color Scheme for Vim" />
      <published>2016-12-23T00:00:00-08:00</published>
      <updated>2016-12-23T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/vim_eldar</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/vim-eldar/"><![CDATA[<p>I use <a href="https://www.vim.org">Vim</a> (actually <a href="https://neovim.io">Neovim</a>) as my primary text editor. I
write Python in it, against all advice I write C++ in it, and I even write
these posts in it. Vim is highly customizable, and because its configuration
files are just plaintext, <a href="https://github.com/agude/dotfiles/tree/master/vim">my environment</a> can be set up with a
simple <code class="language-plaintext highlighter-rouge">git pull</code>. This has made it worth spending time customizing Vim to be
exactly the editor I want; one of those customizations is my color scheme:
<a href="https://github.com/agude/vim-eldar"><strong>Eldar</strong></a>.</p>

<h2 id="eldar">Eldar</h2>

<p>Eldar is based on <a href="https://github.com/vim/vim/blob/master/runtime/colors/elflord.vim">elflord</a>, one of the default Vim color schemes. I
discovered it when I first started using Vim, and I grew to love its 16 bright
colors. But elflord isn’t actually a 16 color scheme, so when I finally fixed
my terminal to use all 256 colors, elflord no longer looked the way I wanted.
I decided to create my own color scheme, one designed to use only 16 colors
from the start. The fact that designing the scheme served double duty as
<a href="https://www.chronicle.com/article/How-to-ProcrastinateStill/93959">structured procrastination</a> to avoid <a href="/blog/my-phd-thesis/">my thesis</a> (also written in
Vim) was merely a bonus!</p>

<h3 id="colors">Colors</h3>

<p>Eldar looks great in both the GUI and terminal because it uses fewer than 16
colors. Eldar uses the <a href="http://tango.freedesktop.org/Tango_Icon_Theme_Guidelines#Color_Palette">Tango color palette</a> in the GUI by default, but
these colors can be overridden by setting <code class="language-plaintext highlighter-rouge">g:eldar_*</code> variables. Eldar
uses both colors and font weight to differentiate various elements. It is not
a subtle scheme; it uses white on black text and there is high contrast
between nine colors. An example of the color scheme for Python with spell
checking turned on and the word ‘<em>Fibonacci</em>’ highlighted:</p>

<p><img src="/files/vim_eldar/eldar_python.png" alt="An example Python file using the Eldar color scheme" /></p>

<h3 id="comparison">Comparison</h3>

<p>Eldar defines many more <code class="language-plaintext highlighter-rouge">highlight</code> groups than elflord, allowing additional
elements to be themed. It colors function declarations and control statements.
Eldar uses different colors for strings and for numbers. It also supports
italic and bold text (for example, in <a href="https://daringfireball.net/projects/markdown/">Markdown</a> and
<a href="https://www.latex-project.org/">LaTeX</a>). Eldar uses underlining for spelling errors just like your
browser or document editor.</p>

<p>A comparison of Eldar to 16 color elflord is shown below. Spell checking is
turned on and the word ‘<em>Handle</em>’ has been highlighted.</p>

<p><img src="/files/vim_eldar/eldar_elflord_cpp.gif" alt="Eldar vs elflord gif" /></p>

<h3 id="installation">Installation</h3>

<p>You can install Eldar by downloading the <a href="https://github.com/agude/vim-eldar/blob/master/colors/eldar.vim"><code class="language-plaintext highlighter-rouge">eldar.vim</code> file</a> and placing
it in <code class="language-plaintext highlighter-rouge">~/.vim/colors/</code>, or by using whatever plugin manager you prefer. For
example, with <a href="https://github.com/junegunn/vim-plug">vim-plug</a>:</p>

<div class="language-vim highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Plug <span class="s1">'agude/vim-eldar'</span>
</code></pre></div></div>

<p>You can activate the color scheme with:</p>

<div class="language-vim highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">:</span><span class="k">colorscheme</span> eldar
</code></pre></div></div>

<p>Or put it in your <code class="language-plaintext highlighter-rouge">.vimrc</code> so it activates every time you open Vim:</p>

<div class="language-vim highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c">"------------------------</span>
<span class="c">" Syntax: highlighting</span>
<span class="c">"------------------------</span>
<span class="k">if</span> <span class="nb">has</span><span class="p">(</span><span class="s1">'syntax'</span><span class="p">)</span>
    <span class="nb">syntax</span> enable             " Turn <span class="k">on</span> <span class="nb">syntax</span> highlighting
    <span class="k">silent</span><span class="p">!</span> <span class="k">colorscheme</span> eldar " Custom color scheme
<span class="k">endif</span>
</code></pre></div></div>

<p>Give Eldar a try! If something doesn’t work, please report it on <a href="https://github.com/agude/vim-eldar">Eldar’s
Github page</a>. Thanks!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="my-projects" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Check out Eldar, my custom Vim color scheme based on elflord. It is a bright, high-contrast theme that looks great in the terminal or GUI!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/vim_eldar/eldar_logo.png" />
        <media:content medium="image" url="https://alexgude.com/files/vim_eldar/eldar_logo.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Lab41 Reading Group: Swapout: Learning an Ensemble of Deep Architectures</title>
      <link href="https://alexgude.com/blog/lab41-swapout/" rel="alternate" type="text/html" title="Lab41 Reading Group: Swapout: Learning an Ensemble of Deep Architectures" />
      <published>2016-12-12T00:00:00-08:00</published>
      <updated>2016-12-12T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/lab41_swapout</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-swapout/"><![CDATA[<p>Next up for the reading group is <a href="https://papers.neurips.cc/paper_files/paper/2016/hash/c51ce410c124a10e0db5e4b97fc2af39-Abstract.html">a paper about a new stochastic training
method</a> written by Saurabh Singh, Derek Hoiem, and David Forsyth of the
University of Illinois at Urbana–Champaign.<sup style="anchor-name:--fnref-singh" id="fnref:singh"><a href="#fn:singh" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Their new training method
is like <a href="https://arxiv.org/abs/1207.0580">dropout</a>,<sup style="anchor-name:--fnref-hinton" id="fnref:hinton"><a href="#fn:hinton" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> <a href="https://doi.org/10.1007/978-3-319-46493-0_39">stochastic depth</a>,<sup style="anchor-name:--fnref-huang" id="fnref:huang"><a href="#fn:huang" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> and
<a href="https://doi.org/10.1109/CVPR.2016.90">ResNets</a><sup style="anchor-name:--fnref-he" id="fnref:he"><a href="#fn:he" class="footnote" rel="footnote" role="doc-noteref">4</a></sup> but with its own special twist. I recommend picking up the
paper after going through this post, it is very readable and includes an
excellent section on performing inference with a stochastically trained
network that I will only touch on.</p>

<p>As you may recall, dropout works by randomly setting individual neuron outputs
in a network to zero, essentially dropping those neurons from training and
hence forcing the network to use a variety of signals instead of over-training
on one. Stochastic depth (<a href="/blog/lab41-stochastic-depth/">covered in a previous post</a>) is similar,
but instead of dropping neurons it bypasses whole layers! We can think of
these operations a little more mathematically, but first I’ll have to define
some notation.</p>

<p>I’ll use block to mean a set of layers in some specific configuration (for
example, a convolution followed by a ReLU), and a unit to be one of the
computational nodes within the block (basically a neuron). \(X\) will be the
input from the previous block, and \(F(X)\) will be the output from a unit
within the current block.</p>

<p>Using this notation then, we can think about ResNets as consisting of blocks
where ever unit in the block always reports \(X + F(X)\). A standard,
feed-forward layer can be viewed in this framework as well, with each unit
always reporting \(F(X)\). The paper includes a figure, which I’ve edited and
included below, showing feed-forward and ResNets in this scheme:</p>

<p><img src="/files/swapout//feed_and_res.svg" alt="A diagram representation of blocks and units for a feed-forward and ResNet
architecture" /></p>

<p>Things become more interesting when we start thinking about stochastic
training methods in this manner. Dropout can be thought of as randomly
selecting the output for each unit from the following set of possible
outcomes: \(\{0, F(X)\}\). Likewise, stochastic depth can be thought of as
randomly selecting between the outcomes \(\{X, F(X)\}\) for each block, so
that every unit in the block returns \(X\) or \(F(X)\) together. Both of these
training methods are shown in the figure below, which is again has been
modified from the paper:</p>

<p><img src="/files/swapout//drop_and_depth.svg" alt="A diagram representation of blocks and units for dropout and stochastic
depth architecture" /></p>

<p>So now that I’ve laid the groundwork, what does swapout add? Well, add isn’t
really the right word, swapout combines! It randomly selects from the four
possible outcomes mentioned above: feed-forward, ResNet, dropout, and
stochastic depth. They do this by allowing each unit to randomly select from
the following outcomes: \(\{0, X, F(X), X + F(X)\}\). Therefore, swapout
samples from every possible stochastic depth and ResNet architecture, both
including and not including dropout!</p>

<p>In addition to swapout, the authors define a simpler version called
skipforward. Skipforward only allows units to select from the outcomes \(\{X,
F(X)\}\), that is limiting the choice to only stochastic depth and
feed-forward. Both of these architectures are shown in the figure below, which
is again from the paper with modification:</p>

<p><img src="/files/swapout//swapout.svg" alt="A diagram representation of blocks and units for skipforward and swapout
depth architecture" /></p>

<p>One of the dilemmas when using stochastic training methods is: how do I use
the network at inference time? When training the network is constantly
mutating as units pick different ways of behaving, but at inference time that
network needs to be roughly static so that the same input will always yield
the same prediction. We can make the network static in two ways:</p>

<ol>
  <li>
    <p><strong>Deterministic inference</strong>: All values are replaced by their expectation
value. That is, the unit that was dropped half the time is set to 50%
weight.</p>
  </li>
  <li>
    <p><strong>Stochastic inference</strong>: Several versions of the network are randomly
generated at the end of training and their results averaged. A unit that is
dropped half the time would (by chance) appear in half randomly generated
versions of the network.</p>
  </li>
</ol>

<p>Although it seems like deterministic inference should be faster (because it
does not require running multiple networks) it has several drawbacks. The
first drawback is that you can not actually calculate the true expectation
value for a swapout network, only approximate it. The second is the fact that
<a href="https://arxiv.org/abs/1502.03167">batch normalization</a><sup style="anchor-name:--fnref-ioffe" id="fnref:ioffe"><a href="#fn:ioffe" class="footnote" rel="footnote" role="doc-noteref">5</a></sup>—one of the most powerful training
methodologies—does not work with deterministic inference. The authors
conclude (through testing) that stochastic inference works best.</p>

<p>The authors test swapout and skipforward networks against networks trained
with stochastic depth, dropout, and various ResNet architectures. They
conclude:</p>

<ul>
  <li>
    <p>Swapout improves results as compared with ResNet</p>
  </li>
  <li>
    <p>Stochastic inference beats deterministic, even with only a few averaged results</p>
  </li>
  <li>
    <p>Increasing the width of a network greatly improves performance</p>
  </li>
</ul>

<p>One final note on the paper: the way they define the various operations as
random selections from a set of possible outcomes is, for me, a very intuitive
way to think about them. I would love to see other papers use a similar
framework for describing their network modifications!</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:singh">

      <p><span class="citation">Singh, Saurabh and Hoiem, Derek and Forsyth, David. <a href="https://papers.neurips.cc/paper_files/paper/2016/hash/c51ce410c124a10e0db5e4b97fc2af39-Abstract.html">“Swapout: learning an ensemble of deep architectures”</a> <cite>Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS’16)</cite>. Curran Associates Inc. 2016. pp. 28–36.</span> <a href="#fnref:singh" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:hinton">

      <p><span class="citation">Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. <a href="https://api.semanticscholar.org/CorpusID:14832074">“Improving neural networks by preventing co-adaptation of feature detectors”</a> <cite>arXiv</cite>. vol. abs/1207.0580. 2012.</span> <a href="#fnref:hinton" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:huang">

      <p><span class="citation">Huang, Gao and Sun, Yu and Liu, Zhuang and Sedra, Daniel and Weinberger, Kilian Q. “Deep Networks with Stochastic Depth” <cite>Computer Vision – ECCV 2016</cite>. Edited by Leibe, Bastian and Matas, Jiri and Sebe, Nicu and Welling, Max. Springer International Publishing. 2016. pp. 646–661. doi: <a href="https://doi.org/10.1007/978-3-319-46493-0_39">10.1007/978-3-319-46493-0_39</a>.</span> <a href="#fnref:huang" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:he">

      <p><span class="citation">He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian. “Deep Residual Learning for Image Recognition” <cite>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</cite>. 2016. pp. 770–778. doi: <a href="https://doi.org/10.1109/CVPR.2016.90">10.1109/CVPR.2016.90</a>.</span> <a href="#fnref:he" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ioffe">

      <p><span class="citation">Ioffe, Sergey and Szegedy, Christian. <a href="https://arxiv.org/abs/1502.03167">“Batch normalization: accelerating deep network training by reducing internal covariate shift”</a> <cite>Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37</cite>. JMLR.org. 2015. pp. 448–456.</span> <a href="#fnref:ioffe" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
          <category term="lab41" />
        
      

      

      
      
        <summary type="html"><![CDATA[Want to train a network but unsure about dropout vs. stochastic depth? Should you use a ResNet? Stop worrying and use Swapout; it does all that and more!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/swapout/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/swapout/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SWITRS: On What Days Do People Crash?</title>
      <link href="https://alexgude.com/blog/switrs-crashes-by-date/" rel="alternate" type="text/html" title="SWITRS: On What Days Do People Crash?" />
      <published>2016-12-02T00:00:00-08:00</published>
      <updated>2016-12-02T00:00:00-08:00</updated>
      <id>https://alexgude.com/blog/switrs_crashes_by_date</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-crashes-by-date/"><![CDATA[<p>The <a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">Statewide Integrated Traffic Records System (SWITRS)</a> contains a
wealth of information, enough to determine who, where, when, and sometimes why
and how for every traffic collision in California. Today, with the assistance
of my <a href="https://github.com/agude/SWITRS-to-SQLite">SWITRS-to-SQLite script</a> (<a href="/blog/switrs-to-sqlite/">discussed previously</a>), I’m
going to look at when car crashes happen, and specifically on what dates.</p>

<p>As always, the Jupyter notebook used to do this analysis can be found
<a href="/files/switrs-accidents-by-date/SWITRS%20Crash%20Dates.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs-accidents-by-date/SWITRS%20Crash%20Dates.ipynb">rendered on Github</a>).</p>

<h2 id="data-selection">Data Selection</h2>

<p>The data was selected using the following query:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span> <span class="n">Collision_Date</span> <span class="k">FROM</span> <span class="n">Collision</span>
<span class="k">WHERE</span> <span class="n">Collision_Date</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="k">NULL</span>
<span class="k">AND</span> <span class="n">Collision_Date</span> <span class="o">&lt;=</span> <span class="s1">'2015-12-31'</span>  <span class="c1">-- 2016 is incomplete</span>
<span class="c1">-- We only want car crashes</span>
<span class="k">AND</span> <span class="n">Bicycle_Collision</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="mi">1</span>
<span class="k">AND</span> <span class="n">Motorcycle_Collision</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="mi">1</span>
<span class="k">AND</span> <span class="n">Pedestrian_Collision</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="mi">1</span>
<span class="k">AND</span> <span class="n">Truck_Collision</span> <span class="k">IS</span> <span class="k">NOT</span> <span class="mi">1</span>

</code></pre></div></div>

<p>This selects every car crash that happened before 2016 that has a collision
date stored. The current year, 2016, is excluded because the data from it is
incomplete.</p>

<h2 id="crashes-per-week">Crashes per Week</h2>

<p>The first thing to look at is crashes as a function of time. Below, I plot
crashes per week to make the trends clearer; plotting per day results in too
many points to separate by eye.</p>

<p><a href="/files/switrs-accidents-by-date/accidents_per_week_in_california.svg"><img src="/files/switrs-accidents-by-date/accidents_per_week_in_california.svg" alt="Line plot showing crashes per week from 2001 to
2015" /></a></p>

<p>The week-to-week variation is rather significant, but two major trends are
obvious:</p>

<ol>
  <li>
    <p>The total number of crashes has been decreasing over the past few years,
with a big drop in 2008, but is now rising sharply in 2015.</p>
  </li>
  <li>
    <p>Each year is similar, with a mid-year lull and wildly varying
increases and decreases right before the end of the year.</p>
  </li>
</ol>

<p>The first trend is easy to explain: the <a href="https://en.wikipedia.org/wiki/Great_Recession">Great Recession</a> put many people
out of work, who then stopped commuting. The second trend is also due to
reduced driving; we’ll look at it in detail below.</p>

<h2 id="day-by-day">Day-by-Day</h2>

<p>To explore the second trend, we’ll need to look at the data day-by-day instead
of a week at a time. Below is a plot of the average number of crashes on
each day of the year. The average is calculated by summing the number of
crashes on a specific day (say, September 22nd) across the years 2001 to
2015. The sum is then divided by the number of times that specific day
appeared in the timespan (15, except for the <a href="https://en.wikipedia.org/wiki/February_29">leap day</a>, which only
appears 3 times).</p>

<p><a href="/files/switrs-accidents-by-date/mean_accidents_by_date.svg"><img src="/files/switrs-accidents-by-date/mean_accidents_by_date.svg" alt="Line plot showing average crashes by day of the
year" /></a></p>

<p>Holidays account for the extrema, with the minimum number of crashes taking
place on Christmas, and the maximum number taking place on Halloween. In fact,
many of the local maxima and minima are also holidays! Some create obvious,
multi-day patterns (like Thanksgiving) because they are floating holidays
while others (like Christmas and Halloween) create massive, single-day dips or
spikes because they happen on the same day every year. Holidays that fewer
people get off from work, like Washington’s Birthday and Columbus Day, show
almost no deviation from the surrounding dates.</p>

<p>But perhaps the most interesting dates are the holidays where the number of
crashes <strong>increases</strong>! Halloween is the most obvious of these, and sets the
record for the <strong>highest number of crashes</strong>, but Valentine’s Day and St.
Patrick’s Day also show increases. I believe there are two reasons these
holidays have higher than normal crash counts. First, these are not
generally paid holidays, so the normal number of commute-related crashes
happen. Second, these holidays are celebrated away from home after work
(for drinks, dates, or candy), and so people drive more on these days than
they would otherwise. I suspect that there is a third reason behind
Halloween’s high crash count: a higher than average number of pedestrians
being out and about leading to a higher than average number of collisions
involving pedestrians. I plan to look at pedestrian incidents in a later blog
post.</p>

<h2 id="day-of-the-week">Day of the Week</h2>

<p>Finally, let’s look at crashes by day of the week. On weekends, like holidays,
we would expect most people to not go to work. Below is a <a href="https://en.wikipedia.org/wiki/Violin_plot">violin
plot</a> of crashes by day of the week. The width of each “violin”
indicates the number of days with that value while the center line indicates
the median, and the two outer lines indicate the interquartile.</p>

<p><a href="/files/switrs-accidents-by-date/accidents_by_day_of_the_week.svg"><img src="/files/switrs-accidents-by-date/accidents_by_day_of_the_week.svg" alt="Violin plot showing crashes by day of the
week" /></a></p>

<p>The distribution for each day of the week is bimodal. This is due to the <a href="#crashes-per-week">two
plateaus in crash rates</a>: a high one from 2001–2006, and a lower one
from 2011–2014. The first four weekdays have roughly the same number of
crashes. Friday has more, presumably because people are more likely to go out
after work. Saturday drops to a level slightly below the weekdays, though not
by much, and Sunday has the lowest crash count.</p>

<p>In the end the results are not too surprising: car crashes happen when people
are driving, not when they’re sitting at home celebrating!</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[What day of the year has the most car crashes? The fewest? Find out as I look at California's crash data! Hint: they're both holidays!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs-accidents-by-date/1923_dc_car_crash.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs-accidents-by-date/1923_dc_car_crash.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Introducing ‘SWITRS to SQLite’</title>
      <link href="https://alexgude.com/blog/switrs-to-sqlite/" rel="alternate" type="text/html" title="Introducing ‘SWITRS to SQLite’" />
      <published>2016-11-01T00:00:00-07:00</published>
      <updated>2016-11-01T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/switrs_to_sqlite</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/switrs-to-sqlite/"><![CDATA[<p>The State of California maintains a database called the <a href="https://www.chp.ca.gov/programs-services/services-information/switrs-statewide-integrated-traffic-records-system/">Statewide Integrated
Traffic Records System (SWITRS)</a>. It contains a record of every
traffic collision that has been reported in the state—the time of the crash,
the location, the vehicles involved, and the reason for the crash. Even
better, it is <a href="https://github.com/agude/SWITRS-to-SQLite/blob/master/requesting_data.md">publicly available</a>!</p>

<p>Unfortunately, the data is delivered as a set of large <a href="https://en.wikipedia.org/wiki/Comma-separated_values">CSV files</a>.
Normally you could just load them into <a href="https://pandas.pydata.org/">Pandas</a>, but there is one, big
problem: the data is spread across three files! This means you must join the
rows between them to select the incidents you are looking for. Pandas can do
these joins, but not without overflowing the memory on my laptop. If only the
data were in a proper database!</p>

<h2 id="switrs-to-sqlite">SWITRS-to-SQLite</h2>

<p>To solve this problem, I wrote <a href="https://github.com/agude/SWITRS-to-SQLite">SWITRS-to-SQLite</a>. SWITRS-to-SQLite is a
Python script that takes the three CSV files returned by SWITRS and converts
them into a <a href="https://sqlite.org/">SQLite3 database</a>. This allows you to perform standard
<a href="https://en.wikipedia.org/wiki/SQL">SQL queries</a> on the data before pulling it into an analysis system like
Pandas. Additionally, the script does some data cleanup like converting the
various null value indicators to a true <code class="language-plaintext highlighter-rouge">NULL</code>, and converting the date and
time information to a form recognized by SQLite.</p>

<h3 id="installation-and-running">Installation and Running</h3>

<p>The best way to install SWITRS-to-SQLite is with
<a href="https://docs.astral.sh/uv/">UV</a>:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>uv tool <span class="nb">install </span>switrs-to-sqlite
</code></pre></div></div>

<p>Or with standard <code class="language-plaintext highlighter-rouge">pip</code>:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>switrs-to-sqlite
</code></pre></div></div>

<p>Running the script on the downloaded data is simple:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code>switrs_to_sqlite <span class="se">\</span>
CollisionRecords.txt <span class="se">\</span>
PartyRecords.txt <span class="se">\</span>
VictimRecords.txt
</code></pre></div></div>

<p>This will run for a while (about a 12 minutes on my ancient desktop) and
produce a SQLite3 file named <code class="language-plaintext highlighter-rouge">switrs.sqlite3</code>.</p>

<h3 id="crash-mapping-example">Crash Mapping Example</h3>

<p>Now that we have the SQLite file, let us make a map of all recorded crashes.
We load the file and select all incidents with GPS coordinates as follows:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="n">pandas</span> <span class="k">as</span> <span class="n">pd</span>
<span class="kn">import</span> <span class="n">sqlite3</span>

<span class="c1"># Read sqlite query results into a pandas DataFrame
</span><span class="k">with</span> <span class="n">sqlite3</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="sh">"</span><span class="s">./switrs.sqlite3</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">con</span><span class="p">:</span>

    <span class="n">query</span> <span class="o">=</span> <span class="p">(</span>
        <span class="sh">"</span><span class="s">SELECT Latitude, Longitude </span><span class="sh">"</span>
        <span class="sh">"</span><span class="s">FROM Collision AS C </span><span class="sh">"</span>
        <span class="sh">"</span><span class="s">WHERE Latitude IS NOT NULL AND Longitude IS NOT NULL</span><span class="sh">"</span>
    <span class="p">)</span>

    <span class="c1"># Construct a Dataframe from the results
</span>    <span class="n">df</span> <span class="o">=</span> <span class="n">pd</span><span class="p">.</span><span class="nf">read_sql_query</span><span class="p">(</span><span class="n">query</span><span class="p">,</span> <span class="n">con</span><span class="p">)</span>
</code></pre></div></div>

<p>Then making a map is simple:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">mpl_toolkits.basemap</span> <span class="kn">import</span> <span class="n">Basemap</span>
<span class="kn">import</span> <span class="n">matplotlib.pyplot</span> <span class="k">as</span> <span class="n">plt</span>

<span class="n">fig</span> <span class="o">=</span> <span class="n">plt</span><span class="p">.</span><span class="nf">figure</span><span class="p">(</span><span class="n">figsize</span><span class="o">=</span><span class="p">(</span><span class="mi">20</span><span class="p">,</span><span class="mi">20</span><span class="p">))</span>

<span class="n">basemap</span> <span class="o">=</span> <span class="nc">Basemap</span><span class="p">(</span>
    <span class="n">projection</span><span class="o">=</span><span class="sh">'</span><span class="s">gall</span><span class="sh">'</span><span class="p">,</span>
    <span class="n">llcrnrlon</span> <span class="o">=</span> <span class="o">-</span><span class="mi">126</span><span class="p">,</span>   <span class="c1"># lower-left corner longitude
</span>    <span class="n">llcrnrlat</span> <span class="o">=</span> <span class="mi">32</span><span class="p">,</span>     <span class="c1"># lower-left corner latitude
</span>    <span class="n">urcrnrlon</span> <span class="o">=</span> <span class="o">-</span><span class="mi">113</span><span class="p">,</span>   <span class="c1"># upper-right corner longitude
</span>    <span class="n">urcrnrlat</span> <span class="o">=</span> <span class="mi">43</span><span class="p">,</span>     <span class="c1"># upper-right corner latitude
</span><span class="p">)</span>

<span class="n">x</span><span class="p">,</span> <span class="n">y</span> <span class="o">=</span> <span class="nf">basemap</span><span class="p">(</span><span class="n">df</span><span class="p">[</span><span class="sh">'</span><span class="s">Longitude</span><span class="sh">'</span><span class="p">].</span><span class="n">values</span><span class="p">,</span> <span class="n">df</span><span class="p">[</span><span class="sh">'</span><span class="s">Latitude</span><span class="sh">'</span><span class="p">].</span><span class="n">values</span><span class="p">)</span>

<span class="n">basemap</span><span class="p">.</span><span class="nf">plot</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="n">y</span><span class="p">,</span> <span class="sh">'</span><span class="s">k.</span><span class="sh">'</span><span class="p">,</span> <span class="n">markersize</span><span class="o">=</span><span class="mf">1.5</span><span class="p">)</span>
</code></pre></div></div>

<p>This gives us a map of the locations of all the crashes in the state of
California from 2001 to 2016:</p>

<p><a href="/files/switrs_to_sqlite/switrs_crash_map.png"><img src="/files/switrs_to_sqlite/switrs_crash_map.png" alt="A map of the location of all the crashes in the state of California from
2001 to 2016" /></a></p>

<p>There are some weird artifacts and grid patterns that show up which are not
due to our mapping but are inherent in the data. Some further clean up will be
necessary before doing any analysis! A Jupyter notebook used to make the map
can be found <a href="/files/switrs_to_sqlite/SWITRS%20Crash%20Map.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/switrs_to_sqlite/SWITRS%20Crash%20Map.ipynb">rendered on Github</a>).</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="california-traffic-data" />
        
          <category term="my-projects" />
        
      

      

      
      
        <summary type="html"><![CDATA[The State of California stores information about all the traffic collisions in the state in the SWITRS database; this script lets you convert it to SQLite for easy querying!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/switrs_to_sqlite/chp.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/switrs_to_sqlite/chp.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Lab41 Reading Group: Skip-Thought Vectors</title>
      <link href="https://alexgude.com/blog/lab41-skipthought-vectors/" rel="alternate" type="text/html" title="Lab41 Reading Group: Skip-Thought Vectors" />
      <published>2016-10-31T00:00:00-07:00</published>
      <updated>2016-10-31T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_skipthought_vectors</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-skipthought-vectors/"><![CDATA[<p>Continuing the tour of older papers that started with our <a href="/blog/lab41-resnet/">ResNet blog
post</a>, we now take on <a href="https://arxiv.org/abs/1506.06726"><strong>Skip-Thought Vectors</strong></a> by <a href="https://scholar.google.com/citations?user=W8zwlYQAAAAJ&amp;hl=en">Kiros</a>
<em>et al.</em><sup style="anchor-name:--fnref-kiros" id="fnref:kiros"><a href="#fn:kiros" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> Their goal was to come up with a useful embedding for
sentences that was not tuned for a single task and did not require labeled
data to train. They took inspiration from Word2Vec skip-gram (you can find <a href="/blog/lab41-python2vec/#word-embeddings">my
explanation of that algorithm here</a>) and attempt to extend it to
sentences.</p>

<p>Skip-thought vectors are created using an encoder-decoder model. The encoder
takes in the training sentence and outputs a vector. There are two decoders
both of which take the vector as input. The first attempts to predict the
previous sentence and the second attempts to predict the next sentence. Both
the encoder and decoder are constructed from recurrent neural networks (RNN).
Multiple encoder types are tried including <strong>uni-skip</strong>, <strong>bi-skip</strong>, and
<strong>combine-skip</strong>. Uni-skip reads the sentence in the forward direction.
Bi-skip reads the sentence forwards and backwards and concatenates the
results. Combined-skip concatenates the vectors from uni- and bi-skip. Only
minimal tokenization is done to the input sentences. A diagram indicating the
input sentence and the two predicted sentences is shown below.</p>

<figure>
  
  <a href="/files/skip-thought//st_example.png">
    <img src="/files/skip-thought//st_example.png" alt="Architecture of the Skip-thought vectors model. A central
  sentence 'I could see the cat on the steps' is processed by an encoder
  (sequence of circles connected by arrows). The final state of the encoder
  (highlighted by a dashed box) is then used by two separate decoders to
  predict the preceding sentence ('I got back home &lt;eos&gt;') and the subsequent
  sentence ('This was strange &lt;eos&gt;'). Decoder states are shown as red circles
  for the previous sentence and green circles for the next." decoding="async" />
  </a>
  
  
  
  <figcaption>Given a sentence (the grey dots), skip-thought attempts to predict
  the preceding sentence (red dots) and the next sentence (green dots). Figure
  from the paper.</figcaption>
  
</figure>

<p>Their model requires groups of sentences in order to train, and so trained on
the <strong>BookCorpus Dataset</strong>. The dataset consists of novels by unpublished authors
and is (unsurprisingly) dominated by romance and fantasy novels. This “bias”
in the dataset will become apparent later when discussing some of the
sentences used to test the skip-thought model; some of the retrieved sentences
are quite exciting!</p>

<p>Building a model that accounts for the meaning of an entire sentence is tough
because language is remarkably flexible. Changing a single word can either
completely change the meaning of a sentence or leave it unaltered. The same is
true for moving words around. As an example:</p>

<blockquote>
  <p>One <strong>difficulty</strong> in building a model to handle sentences is that a single
word can be changed and yet the meaning of the sentence is the same.</p>
</blockquote>

<p>Put a different way:</p>

<blockquote>
  <p>One <strong>challenge</strong> in building a model to handle sentences is that a single
word can be changed and yet the meaning of the sentence is the same.</p>
</blockquote>

<p>Changing a single word has had almost no effect on the meaning of that
sentence. To account for these word level changes, the skip-thought model
needs to be able to handle a large variety of words, some of which were not
present in the training sentences. The authors solve this by using a
pre-trained continuous bag-of-words (CBOW) Word2Vec model and learning a
translation from the Word2Vec vectors to the word vectors in their sentences.
Below are shown the nearest neighbor words after the vocabulary expansion
using query words that do not appear in the training vocabulary:</p>

<figure>
  
  <a href="/files/skip-thought//words.png">
    <img src="/files/skip-thought//words.png" alt="Table illustrating semantically similar words generated by the
  Skip-thought vectors model (Kiros et al.). Column headers are target words:
  'choreograph', 'modulation', 'vindicate', 'neuronal', 'screwy', 'Mykonos',
  and 'Tupac'. Each column lists related terms, e.g., under 'choreograph' are
  'choreography', 'rehearse'; under 'Mykonos' are 'Glyfada', 'Santorini';
  under 'Tupac' are '2Pac', 'Cormega'." decoding="async" />
  </a>
  
  
  
  <figcaption>Nearest neighbor words for various words that were not included in
  the training vocabulary. Table from the paper.</figcaption>
  
</figure>

<p>So how well does the model work? One way to probe it is to retrieve the
closest sentence to a query sentence; here are some examples:</p>

<blockquote>
  <p><strong>Query</strong>: “I’m sure you’ll have a glamorous evening,” she said, giving an
exaggerated wink.</p>
</blockquote>

<blockquote>
  <p><strong>Retrieved</strong>: “I’m really glad you came to the party tonight,” he said,
turning to her.</p>
</blockquote>

<p>And:</p>

<blockquote>
  <p><strong>Query</strong>: Although she could tell he hadn’t been too interested in any of
their other chitchat, he seemed genuinely curious about this.</p>
</blockquote>

<blockquote>
  <p><strong>Retrieved</strong>: Although he hadn’t been following her career with a
microscope, he’d definitely taken notice of her appearance.</p>
</blockquote>

<p>The sentences are in fact very similar in both structure and meaning (and a
bit salacious, as I warned earlier) so the model appears to be doing a good
job.</p>

<p>To perform more rigorous experimentation, and to test the value of
skip-thought vectors as a generic sentence feature extractor, the authors run
the model through a series of tasks using the encoded vectors with simple,
linear classifiers trained on top of them.</p>

<p>They find that their generic skip-thought representation performs very well
for detecting the semantic relatedness of two sentences and for detecting
where a sentence is paraphrasing another one. Skip-thought vectors perform
relatively well for image retrieval and captioning (where they use
<a href="https://arxiv.org/abs/1409.1556">VGG</a><sup style="anchor-name:--fnref-simonyan" id="fnref:simonyan"><a href="#fn:simonyan" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> to extract image feature vectors). Skip-thought performs
poorly for sentiment analysis, producing equivalent results to various bag of
word models but at a much higher computational cost.</p>

<p>We have used skip-thought vectors a little bit at the Lab, most recently for
the <a href="https://medium.com/gab41/tell-me-something-i-dont-know-detecting-novelty-and-redundancy-with-natural-language-processing-818124e4013c">Pythia challenge</a>. We found them to be useful for novelty
detection, but incredibly slow. Running skip-thought vectors on a corpus of
about 20,000 documents took many hours, whereas simpler (and as effective)
methods took seconds or minutes.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:kiros">

      <p><span class="citation">Kiros, Ryan, <abbr class="etal">et al.</abbr> <a href="https://arxiv.org/abs/1506.06726">“Skip-thought vectors”</a> <cite>Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 2 (NIPS’15)</cite>. MIT Press. 2015. pp. 3294–3302.</span> <a href="#fnref:kiros" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:simonyan">

      <p><span class="citation">Simonyan, K and Zisserman, A. <a href="https://arxiv.org/abs/1409.1556">“Very deep convolutional networks for large-scale image recognition”</a> <cite>3rd International Conference on Learning Representations (ICLR 2015)</cite>. Computational and Biological Learning Society. 2015. pp. 1–14.</span> <a href="#fnref:simonyan" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
          <category term="lab41" />
        
      

      

      
      
        <summary type="html"><![CDATA[Word embeddings are great and should be your first stop for doing word based NLP. But what about sentences? Read on to learn about skip-thought vectors, a sentence embedding algorithm!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/skip-thought/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/skip-thought/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Jupyter Notebooks: Not for Development</title>
      <link href="https://alexgude.com/blog/jupyter-not-for-development/" rel="alternate" type="text/html" title="Jupyter Notebooks: Not for Development" />
      <published>2016-10-17T00:00:00-07:00</published>
      <updated>2016-10-17T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/jupyter_not_for_development</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/jupyter-not-for-development/"><![CDATA[<p><a href="https://jupyter.org/">Jupyter Notebooks</a> are great! They make it really convenient to tinker
with a new library and are excellent for documenting projects that include
code. What Jupyter Notebooks are not great for, and what I find many people
(including my lab) using them for, is development. There are three reasons for
this:</p>

<h3 id="1-version-control">1. Version Control</h3>

<p>Version controlling notebooks is a mess; in addition to code, they also
contain data and output, which results in a high number of changes every time
the notebook is run. Worse, even if the output is identical, things like cell
numbering update every run and so flag the notebook as changed to the version
control system. Further, JSON is already hard to <code class="language-plaintext highlighter-rouge">diff</code>, and adding these
superfluous changes makes it harder still.</p>

<h3 id="2-modularity">2. Modularity</h3>

<p>Other code can not easily call code defined in notebooks. This leads to lots
of duplicated code, and means that notebooks need to either appear at the end
of the pipeline or write to disk to pass on data. This lack of modularity also
makes it difficult to write unit tests to verify the correctness of notebook
code.</p>

<h3 id="3-complex-history">3. Complex History</h3>

<p>Notebooks have complicated history; they cache the results of previous cells
including set variables. Notebooks are so flexible that you will often add and
delete cells when working on them, leaving you in a state with impossible to
remember history. This means that unless you have run the notebook from a
fresh kernel it is possible that the results are dependent on now deleted
cells.</p>

<h2 id="good-uses">Good Uses</h2>

<p>This is not to say that the use of <a href="https://en.wikipedia.org/wiki/Considered_harmful">Jupyter Notebooks should be considered
harmful</a>; they are great for:</p>

<ul>
  <li>
    <p><strong>Exploring data:</strong> In-line plots make it very easy to check
something, make a tweak, and check again. The caching of variables means
that expensive operations can often be called once and the results used for
plot after plot.</p>
  </li>
  <li>
    <p><strong>Testing out new libraries:</strong> Notebooks really shorten the time between
hitting an error, editing, and rerunning code, making them ideal for
trying out new libraries, classes, and functions.</p>
  </li>
  <li>
    <p><strong>Providing a final deliverable:</strong> Notebooks can include runnable code,
text, and images so they make an excellent way to document an analysis and
provide a way for others to interface with it.</p>
  </li>
</ul>

<p>So in closing: I use Jupyter Notebooks where they excel—like documenting
analyses for this blog or tweaking algorithms for
<a href="/blog/whereto-photo/">WhereTo.Photo</a>—and try to stick to pure code for other cases.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="jupyter" />
        
          <category term="software-development" />
        
      

      

      
      
        <summary type="html"><![CDATA[Jupyter Notebooks are great for a lot of things; development of code is not one of them.]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/jupyter_dev/red_spot.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/jupyter_dev/red_spot.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">WhereTo.Photo: Using Data Science to Take Great Photos</title>
      <link href="https://alexgude.com/blog/whereto-photo/" rel="alternate" type="text/html" title="WhereTo.Photo: Using Data Science to Take Great Photos" />
      <published>2016-09-22T00:00:00-07:00</published>
      <updated>2016-09-22T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/whereto_photo</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/whereto-photo/"><![CDATA[<p>After graduating from the University of Minnesota, I moved back to California
to attend <a href="https://web.archive.org/web/20200129121516/https://www.insightdatascience.com/">Insight Data Science</a>. Insight is a seven week program
that takes newly minted PhDs in quantitative fields and grooms them for
careers in data science (I later wrote about <a href="/blog/should-i-go-to-insight/">my experience and whether you
should attend Insight</a>). The first four weeks of the
program focus on building a data product using publicly available data. The
project I built, <strong>Whereto.photo</strong>, tried to answer the question: <em>Where is
the best place in this city to take a picture?</em></p>

<p>In this post I’m going to walk through how I built my project from
brainstorming and data processing to hosting.</p>

<h2 id="project-ideas">Project Ideas</h2>

<p>I had a few ideas about the sort of project I wanted to build when I arrived
at Insight. The ideas were mostly based on my hobbies: cycling, running, and
photography. But I knew that one of the hardest things about making a project
was finding public data to use, so before settling on any idea I checked to
see what data was available. There was not much data on running or
cycling—Strava keeps their site locked down pretty tight—but there were a
lot of photos available from Instagram and Flickr, so I decided to do a photo
project.</p>

<p>When I arrive in a new city I generally already have an idea of what landmarks
I want to shoot, maybe a sunset or a bridge or a <a href="https://en.wikipedia.org/wiki/Gum_Wall">wall caked in gum</a>,
but I don’t know the best place to go to shoot them. So the question I decided
to answer was: <em>Where can I go around here to take the best picture of X?</em></p>

<p>Answering this question would require me to determine three things about the
photos that made up my data:</p>

<ol>
  <li>
    <p>Location</p>
  </li>
  <li>
    <p>Subject</p>
  </li>
  <li>
    <p>Quality</p>
  </li>
</ol>

<h2 id="dataset">Dataset</h2>

<p>I initially explored using data from Instagram, but found that their API made
it difficult to request photos from a single area, which made answering the
location part of my question difficult. I settled on using Flickr because
its API allowed me to ask for all photos within some radius of a fixed point
which made it very easy to download just the photos taken in the cities I was
interested in.</p>

<p>I ended up downloading every photo that was taken between June 1st, 2014 and
May 31st, 2015 within 20km of the centers of San Francisco (Lat 37.74, Lon
-122.42), Seattle (47.61, -122.34), and New York (40.70, -73.98). I included
only public photos that were marked by Flickr as being “safe for work”
(although that tag is only as good as the reviewers on the site). San
Francisco had 255,232 photos, Seattle had 118,464, and New York had 474,649.</p>

<p>Here is every photo in my dataset for San Francisco:</p>

<p><a href="/files/whereto_photo/sf_all_photos.png"><img src="/files/whereto_photo/sf_all_photos.png" alt="Every photo from San Francisco in the dataset" /></a></p>

<h2 id="photo-subject-and-quality">Photo Subject and Quality</h2>

<p>Once I had the data downloaded, the next step was to determine what was in
each photo, and how good the photo was. Unfortunately, computer vision is hard
and not something that I could tackle on the time scale of several weeks. That
left me with only user applied tags as the means of determining the content of
the images.</p>

<p>There are two types of tags applied to Flickr images. The first type are user
applied tags, and the second type are machine applied tags. Machine applied
tags generally just duplicate the <a href="https://en.wikipedia.org/wiki/Exif">EXIF data</a> and so I removed them.
User tags are (as their name suggests) applied by the user and may contain any
information they deem appropriate. These tags often contain subjects (“golden
gate”), locations (“san francisco”), equipment notes (“canon ef 50mm f/1.8”),
and emotions (“sad”). All tags were converted to lowercase to remove
duplicates that differed only by capitalization. Unfortunately
<a href="https://stuvel.eu/flickrapi-doc/">flickrapi</a>, which I used to make my requests, takes multi-word tags and
removes the spaces leaving a single monster word. This meant that when users
accessed my project I also had to strip the spaces from their queries in order
to match tags.</p>

<p>To determine the quality of a photo I had two pieces of metadata I could use:
views and favorites. I decided to use views because favorites were rare; any
click was counted as a view but only logged in users could mark a photo as a
favorite. I would have liked to do comparison testing of views to favorites,
but views worked well enough and I felt my time was better spent elsewhere. To
select the very best photos for each tag, I took the top 10% of photos in
terms of views (or the first 20, whichever was larger) and used that as my
dataset.</p>

<h2 id="the-best-spot">The Best Spot</h2>

<p>Having selected the best photos for each tag, I needed to determine where the
best spot to take a photo of the tagged subject.</p>

<p>My first attempt was to cluster the photos in space, weighted by their
quality, so that the cluster centers would indicate areas of high quality
photos. I tried <a href="https://en.wikipedia.org/wiki/K-means_clustering"><em>k</em>-means clustering</a> and <a href="https://en.wikipedia.org/wiki/DBSCAN">DBSCAN</a> but found
them too inflexible; some tags had a single cluster of photos while others had
many, and the density and spacing between clusters varied too much from tag to
tag for any single set of parameters to work.</p>

<p>My second attempt used <a href="https://en.wikipedia.org/wiki/Kernel_density_estimation">kernel density estimation (KDE)</a> to estimate the
probability density of good photos in the city as a function of location. The
maximum in this distribution was then the “best spot” to take a photo of the
tagged thing. This algorithm worked well, except it favored areas with lots of
photos, not necessarily areas with the best photos. For example, in San
Francisco the maximum of the “flowers” KDE was in the financial district, not
because the best flower photos were taken there, but because every tourist
took dozens of mediocre photos there tagged flowers.</p>

<p>My solution to this was to calculate a second, global KDE for each city using
every photo. I then divided the KDE for a specific tag by the global KDE which
gave me a ratio: <code class="language-plaintext highlighter-rouge">Number of Good Photos / All Photos</code>. The maximum of this
normalized function would be in a location where many good photos with the
specific tag were taken, but few photos in general were taken; essentially the
algorithm now preferred areas with a surprisingly large amount of quality
photos instead of areas with just a large number of photos.</p>

<p>The maximum of this normalized KDE was computed by using the
<a href="https://en.wikipedia.org/wiki/Broyden%E2%80%93Fletcher%E2%80%93Goldfarb%E2%80%93Shanno_algorithm">Broyden–Fletcher–Goldfarb–Shanno algorithm</a> from SciPy. The fitter
often wandered off the edge of the map and into the water, and so a penalty
was applied to all points in the water (using the <code class="language-plaintext highlighter-rouge">basemap.isWater()</code> method).
The starting locations of the minimizer were hand selected to cover the major
land masses in each city.</p>

<p>An example fit for the tag <code class="language-plaintext highlighter-rouge">goldengate</code> is shown below. The blue point is the
estimated best location, the green triangles are the hand selected start
points, the red heat map is the normalized KDE, and black dots are the photos
used in the calculation.</p>

<p><a href="/files/whereto_photo/Goldengate_map.png"><img src="/files/whereto_photo/Goldengate_map.png" alt="The KDE and best photo location for the tag `goldengate`" /></a></p>

<h2 id="the-website">The Website</h2>

<p>The results of the analysis was made available on <strong>Whereto.photo</strong>. The
website was served with <a href="http://flask.pocoo.org/">Flask</a>, and the maxima for
each tag and the associated photos were stored in a MySQL database. The user
input was lowercased and concatenated into a single string and only exact
matches to tags were used.</p>

<p>Here is what the website would show if the user searched for “Golden Gate”:</p>

<p><a href="/files/whereto_photo/Goldengate.png"><img src="/files/whereto_photo/Goldengate.png" alt="The result page of the query &quot;Golden Gate&quot;" /></a></p>

<p>The blue circles are photos, and the blue marker is the predicted best
location. You can see the best location matches the maximum found on the
heat map example above. The user can click on the various photos and a preview
of them would load from Flickr. This allowed the user to verify that the
photos in the predicted best location were in fact great photos. Here is an
example of the user clicking on a photo near the best location for the “Golden
Gate” query:</p>

<p><a href="/files/whereto_photo/Goldengate_with_pic.png"><img src="/files/whereto_photo/Goldengate_with_pic.png" alt="A preview of one of the photos from the query &quot;Golden Gate&quot;" /></a></p>

<p>Sometimes though, the algorithm failed, as in this case for the search term
“Cars”:</p>

<p><a href="/files/whereto_photo/Cars_failure.png"><img src="/files/whereto_photo/Cars_failure.png" alt="A failed search for &quot;Cars&quot;" /></a></p>

<p>In this case the normalization by the global KDE made Treasure Island the
predicted best location because the only photos taken there were for a car
show, giving a very high ratio of good photos. While there were good car
photos taken there, the event was a one time deal and so the recommendation is
not generally useful.</p>

<p>Finally, I’ll leave you with a video of the site to give you a feel for how it
worked:</p>

<!-- WhereTo.Photo Example Youtube Video -->
<style>.embed-container { position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden; max-width: 100%; } .embed-container iframe, .embed-container object, .embed-container embed { position: absolute; top: 0; left: 0; width: 100%; height: 100%; }</style>
<div class="embed-container"><iframe src="https://www.youtube.com/embed/RwkNma7sy2o" frameborder="0" allowfullscreen=""></iframe></div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="my-projects" />
        
          <category term="data-science" />
        
      

      

      
      
        <summary type="html"><![CDATA[Where is the best spot to take a photo in San Francisco? Learn how I answered this question with my Insight Data Science project!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/whereto_photo/WhereTo.Photo.png" />
        <media:content medium="image" url="https://alexgude.com/files/whereto_photo/WhereTo.Photo.png" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Lab41 Reading Group: Deep Residual Learning for Image Recognition</title>
      <link href="https://alexgude.com/blog/lab41-resnet/" rel="alternate" type="text/html" title="Lab41 Reading Group: Deep Residual Learning for Image Recognition" />
      <published>2016-09-08T00:00:00-07:00</published>
      <updated>2016-09-08T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_resnet</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-resnet/"><![CDATA[<p>Today’s <a href="https://arxiv.org/abs/1512.03385">paper offers a new architecture for Convolution Networks</a>. It
was written by He, Zhang, Ren, and Sun from Microsoft Research.<sup style="anchor-name:--fnref-he" id="fnref:he"><a href="#fn:he" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> I’ll warn
you before we start: this paper is ancient. It was published in the dark ages
of deep learning sometime at the end of 2015, which I’m pretty sure means its
original format was papyrus; thankfully someone scanned it so that future
generations could read it. But it is still worth blowing off the dust and
flipping through it because the architecture it proposes has been used time
and time again, including in <a href="/blog/lab41-stochastic-depth/">some of the papers we have previously
read</a>: <a href="https://arxiv.org/abs/1603.09382">Deep Networks with Stochastic Depth</a></p>

<p>He <em>et al.</em> begin by noting a seemingly paradoxical situation: very deep
networks perform more poorly than moderately deep networks, that is, that
while adding layers to a network generally improves the performance, after
some point the new layers begin to hinder the network. They refer to this
effect as network <strong>degradation</strong>.</p>

<p>If you have been following our <a href="/blog/lab41-stochastic-depth/">previous posts</a> this won’t surprise
you; training issues like vanishing gradients become worse as networks get
deeper so you would expect more layers to make the network worse after some
point. But the authors anticipate this line of reasoning and state that
several other deep learning methods, like <a href="https://arxiv.org/abs/1502.03167">batch
normalization</a><sup style="anchor-name:--fnref-ioffe" id="fnref:ioffe"><a href="#fn:ioffe" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> (see <a href="https://medium.com/gab41/batch-normalization-what-the-hey-d480039a9e3b">our post for a summary</a>),
essentially have solved these training issues, and yet the networks still
perform increasingly poorly as their depth increases. For example, they
compare 20- and 56-layer networks and find the 56-layer network performs far
worse; see the image below from their paper.</p>

<figure>
  
  <a href=" https://alexgude.com/files/resnet/20_vs_56.png ">
    <img src=" https://alexgude.com/files/resnet/20_vs_56.png " alt="A plot showing how deeper networks train less well than shallow
  networks." decoding="async" />
  </a>
  
  
  
  <figcaption>Comparison of 20- and 56-layer networks on CIFAR-10. Note that the
  56-layer network performs more poorly in both training and testing.</figcaption>
  
</figure>

<p>The authors then set up a thought experiment (or <a href="https://en.wiktionary.org/wiki/gedankenexperiment">gedankenexperiment</a> if
you’re a recovering physicist like me) to demonstrate that deeper networks
should always perform better. Their argument is as follows:</p>

<ul>
  <li>
    <p>Start with a network that performs well;</p>
  </li>
  <li>
    <p>Add additional layers that are forced to be the identity function, that is,
they simply pass along whatever information arrives at them without change;</p>
  </li>
  <li>
    <p>This network is deeper, but must have the same performance as the original
network by construction since the new layers do not do anything;</p>
  </li>
  <li>
    <p>Layers in a network can learn the identity function, so they should be able
to exactly replicate the performance of this deep network if it is optimal.</p>
  </li>
</ul>

<p>This thought experiment leads them to propose their <strong>deep residual learning</strong>
architecture. They construct their network of what they call residual building
blocks. The image below shows one such block. These blocks have become known
as ResBlocks.</p>

<figure>
  
  <a href=" https://alexgude.com/files/resnet/resblock.svg ">
    <img src=" https://alexgude.com/files/resnet/resblock.svg " alt="A diagram of a ResNet block, or ResBlock." decoding="async" />
  </a>
  
  
  
  <figcaption>A ResBlock; a residual function f(x) is learned on the top and
  information is passed along the bottom unchanged. Image modified from Huang
  <em>et al.</em>'s Stochastic Depth paper.</figcaption>
  
</figure>

<p>The ResBlock is constructed out of normal network layers connected with
<a href="https://en.wikipedia.org/wiki/Rectifier_%28neural_networks%29">rectified linear units</a> (ReLUs) and a pass-through below that feeds
through the information from previous layers unchanged. The network part of
the ResBlock can consist of an arbitrary number of layers, but the simplest is
two.</p>

<p>To get a little into the math behind the ResBlock: let us assume that a set of
layers would perform best if they learned a specific function, \(h(x)\). The
authors note that the residual, \(f(x) = h(x) − x\), can be learned instead
and combined with the original input such that we recover \(h(x)\) as follows:
\(h(x) = f(x) + x\). This can be accomplished by adding a \(+x\) component to
the network, which, thinking back to our thought experiment, is simply the
identity function. The authors hope that adding this “pass-through” to their
layers will aid in training. As with most deep learning, there is only this
intuition backing up the method and not any deeper understanding. However, as
the authors show, it works, and in that end that’s the only thing many of us
practitioners care about.</p>

<p>The paper also explores a few modifications to the ResBlock. The first is
creating bottleneck blocks with three layers where the middle layer constricts
the information flow by using fewer inputs and outputs. The second is testing
different types of pass-through connections including learning a full
projection matrix. Although the more complicated pass-throughs perform better,
they do so only slightly and at the cost of training time.</p>

<p>The rest of the paper tests the performance of the network. The authors find
that their networks perform better than identical networks without the
pass-through; see the image below for their plot showing this. They also find
that they can train far deeper networks and still show improved performance,
culminating in training a 152-layer ResNet that outperforms shallower
networks. They even train a 1202-layer network to prove that it is feasible,
but find that its performance is worse than the other networks examined in the
paper.</p>

<figure>
  
  <a href=" https://alexgude.com/files/resnet/resnet_18_vs_resnet_34.png ">
    <img src=" https://alexgude.com/files/resnet/resnet_18_vs_resnet_34.png " alt="A plot showing how deeper networks train better than shallow
  networks when using ResBlocks." decoding="async" />
  </a>
  
  
  
  <figcaption>A comparison of the performance of two networks: the ones on the
  left do not use ResBlocks, while the ones on the right do. Notice that the
  34-layer network performs better than the 18-layer network, but only when
  using ResBlocks.</figcaption>
  
</figure>

<p>So that’s it! He <em>et al.</em> proposed a new architecture motivated by thought
experiments and the hope that it will work better than previous ones. They
construct several networks, including a few very deep ones, and find that
their new architecture does indeed improve performance of the networks.
Although we don’t gain any further understanding of the underlying principles
of deep learning, we do get a new method of making our networks work better,
and in the end maybe that’s good enough.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:he">

      <p><span class="citation">He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian. “Deep Residual Learning for Image Recognition” <cite>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</cite>. 2016. pp. 770–778. doi: <a href="https://doi.org/10.1109/CVPR.2016.90">10.1109/CVPR.2016.90</a>.</span> <a href="#fnref:he" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ioffe">

      <p><span class="citation">Ioffe, Sergey and Szegedy, Christian. <a href="https://arxiv.org/abs/1502.03167">“Batch normalization: accelerating deep network training by reducing internal covariate shift”</a> <cite>Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37</cite>. JMLR.org. 2015. pp. 448–456.</span> <a href="#fnref:ioffe" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
          <category term="lab41" />
        
      

      

      
      
        <summary type="html"><![CDATA[Inception, AlexNet, VGG... There are so many network architectures, which one should you be using? The one everyone else is: ResNet! Come find out how it works!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/resnet/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/resnet/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Lab41 Reading Group: Deep Compression</title>
      <link href="https://alexgude.com/blog/lab41-deep-compression/" rel="alternate" type="text/html" title="Lab41 Reading Group: Deep Compression" />
      <published>2016-08-09T00:00:00-07:00</published>
      <updated>2016-08-09T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_deep_compression</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-deep-compression/"><![CDATA[<p><a href="https://arxiv.org/abs/1510.00149">The next paper from our reading group</a> is by Song Han, Huizi Mao, and
William J. Dally.<sup style="anchor-name:--fnref-han" id="fnref:han"><a href="#fn:han" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> It won the best paper award at ICLR 2016. It details
three methods of compressing a neural network in order to reduce the size of
the network on disk, improve performance, and decrease run time.</p>

<p>Pre-trained convolutional neural networks are too large for mobile devices:
<a href="http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks">AlexNet</a><sup style="anchor-name:--fnref-krizhevsky" id="fnref:krizhevsky"><a href="#fn:krizhevsky" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> is 240 MB and <a href="https://arxiv.org/abs/1409.1556">VGG-16</a><sup style="anchor-name:--fnref-simonyan" id="fnref:simonyan"><a href="#fn:simonyan" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> is over 552
MB. This seems small when compared to a music library or large video, but the
difference is that the networks reside in memory when running. On mobile
devices SRAM is scarce and DRAM is expensive to access in terms of energy
used. For reference, the authors estimate that a 1 billion node network
running at 20 FPS on your phone would draw nearly 13 Watts from just the DRAM
access alone. If these networks are going to run on a mobile device (where
they could, for example, automatically tag pictures as they are taken) they
must be compressed in some manner. In this paper the authors apply three
compression methods to the weights of various networks and measure the
results. A diagram of the three methods and their results are below; I’ll walk
you through them in more depth in the next few paragraphs.</p>

<figure>
  
  <a href=" /files/deep-compression//compression_stages.png ">
    <img src=" /files/deep-compression//compression_stages.png " alt="A chart summarizing the stages of deep compression." decoding="async" />
  </a>
  
  
  
  <figcaption>A summary of the three stages in the compression pipeline proposed
  by Han <em>et al.</em> Note that the size reduction is cumulative. Image
  from their paper.</figcaption>
  
</figure>

<p>The first compression method is <strong>Network Pruning</strong>. In this method a network
is fully trained and then any connections with a weight below a certain
threshold are removed leaving a sparse network. The sparse network is then
retrained to ensure the remaining connections are used optimally. This form of
compression reduced the size of AlexNet by a factor of 9, and VGG-16 by a
factor of 13. The authors also use a clever data structure that makes use of
variably sized integers to store the network after this compression.</p>

<p>The second compression method is <strong>Trained Quantization and Weight Sharing</strong>.
Here the weights in a network are clustered together with other weights of
similar magnitude, and all these weights are then represented by a single
shared value. The authors use k-means clustering to group weights for sharing.
They explore multiple methods of setting the k centroids and find that a
simple linear spacing of the centroids along the full distribution of weight
values performs best. This compression method reduces the size of the networks
by a factor of 3 or 4. A diagram with an example of this compression technique
is shown below.</p>

<figure>
  
  <a href=" /files/deep-compression//quantization.png ">
    <img src=" /files/deep-compression//quantization.png " alt="A diagram showing how weights are quantized." decoding="async" />
  </a>
  
  
  
  <figcaption>A toy example of trained quantization and weight sharing. On the
  top row, weights of the same color have been clustered and will be replaced
  by a centroid value. On the bottom row, gradients are calculated and used to
  update the centroids. From Han <em>et al.</em></figcaption>
  
</figure>

<p>The third and final compression method is <strong>Huffman Coding</strong>. Huffman coding
is a standard lossless compression technique. The general idea is that it uses
fewer bits to represent data that appears frequently and more bits to
represent data that appears infrequently. For more details see the <a href="https://en.wikipedia.org/wiki/Huffman_coding">Wikipedia
Article</a>. Huffman coding reduces network size by 20% to 30%.</p>

<p>Using all three compression methods leads to a compression factor of 35 times
for AlexNet, and 49 times for VGG-16! This reduces AlexNet to 6.9 MB, and
VGG-16 to under 11.3 MB! Unsurprisingly it is the fully connected layers that
are the largest (90% of the model size), but they also compress the best (96%
of weights pruned in VGG-16). The new, smaller convolutional layers run faster
than their old versions (4 times faster on mobile GPU) and use less energy (4
times less). These results are achieved with no loss in performance! A plot
showing the energy efficiency and speedups due to compression are shown below:</p>

<figure>
  
  <a href=" /files/deep-compression//energy_usage_of_deep_learning.png ">
    <img src=" /files/deep-compression//energy_usage_of_deep_learning.png " alt="A diagram showing energy efficiency and speedups due to compression." decoding="async" />
  </a>
  
  
  
  <figcaption>The energy efficiency and speedups due to compression for various
  layers in the neural networks. The dense bars are the results before
  compression, and the pruned bars are the results after. Note the Y axis is
  log10! From Han, Mao, and Dally's paper.</figcaption>
  
</figure>

<p>Han, Mao, and Dally’s compression techniques achieve an almost perfect result:
the in memory size of a network is reduced, the run speed is increased, and
the energy used to perform the calculation is decreased. Although designed
with mobile in mind, their compression is so successful I would not be
surprised to see it widely supported by the various deep learning frameworks
in the near future.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:han">

      <p><span class="citation">Han, Song and Mao, Huizi and Dally, William J. <a href="https://arxiv.org/abs/1510.00149">“Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding”</a> <cite>arXiv preprint</cite>. October 1, 2015.</span> <a href="#fnref:han" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:krizhevsky">

      <p><span class="citation">Krizhevsky, Alex and Sutskever, Ilya and Hinton, Geoffrey E. <a href="https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf">“ImageNet Classification with Deep Convolutional Neural Networks”</a> <cite>Advances in Neural Information Processing Systems</cite>. Edited by F. Pereira and C.J. Burges and L. Bottou and K.Q. Weinberger. vol. 25. Curran Associates, Inc. 2012.</span> <a href="#fnref:krizhevsky" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:simonyan">

      <p><span class="citation">Simonyan, K and Zisserman, A. <a href="https://arxiv.org/abs/1409.1556">“Very deep convolutional networks for large-scale image recognition”</a> <cite>3rd International Conference on Learning Representations (ICLR 2015)</cite>. Computational and Biological Learning Society. 2015. pp. 1–14.</span> <a href="#fnref:simonyan" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
          <category term="lab41" />
        
      

      

      
      
        <summary type="html"><![CDATA[Deep learning is the future, but how can I fit a battery-drain, half-gigabyte network on my phone? You compress it! Come find out how deep compression saves space and power!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/deep-compression/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/deep-compression/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">SAT2Vec: Word2Vec Versus SAT Analogies</title>
      <link href="https://alexgude.com/blog/sat2vec/" rel="alternate" type="text/html" title="SAT2Vec: Word2Vec Versus SAT Analogies" />
      <published>2016-07-11T00:00:00-07:00</published>
      <updated>2016-07-11T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/SAT2Vec</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/sat2vec/"><![CDATA[<p>Word embeddings, like <a href="https://papers.nips.cc/paper/5021-distributed-representations-of-words-and-phrases-and-their-compositionality.pdf">Word2Vec</a> and <a href="https://nlp.stanford.edu/projects/glove/">GloVe</a>, have
proved to be a powerful way of representing text for machine learning
algorithms. The idea behind these methods is relatively simple: words that are
close to each other in the training text should be close to each other in the
vector space. Of course, you could achieve this by having all the words in the
exact same spot, but that wouldn’t form a useful model, so there is a second
requirement: words that are not close to each other in the text should not be
close to each other in the vector space.</p>

<h2 id="analogies">Analogies</h2>

<p>This simple algorithm produces some neat features, the coolest of which is the
existence of semantic meaning of the directions in the vector space. The
canonical example of this is that the analogy <code class="language-plaintext highlighter-rouge">King : Man :: Queen : Woman</code>
holds true <em>mathematically</em> in the vector space as follows:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>King − Man = Queen − Woman
</code></pre></div></div>

<p>Which is more often rewritten as:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>King − Man + Woman = Queen
</code></pre></div></div>

<p>This shows that one of the directions in our model is “gender”! These
analogies exist for other concepts as well, for example:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Paris − France + Japan = Tokyo
</code></pre></div></div>

<p>And (<a href="https://medium.com/gab41/street-style-guide-vector-transformations-betta-work-2ad8d9829587">as shown by</a> my friend and colleague Patrick Callier):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Workin − Working + Going = Goin
</code></pre></div></div>

<h2 id="word2vec-is-to-the-sat-as">Word2Vec is to the SAT as?</h2>

<p>So with all these analogies embedded in the model, I started thinking back to
when analogies were most prevalent in my life: SAT college entrance exam
preparation! How would the model fare if asked to complete a few SAT
analogies?</p>

<p>To find out, I grabbed Google’s <a href="https://drive.google.com/file/d/0B7XkCwpI5KDYNlNUTTlSS21pQmM/edit?usp=sharing">pretrained Word2Vec model</a>, which
was trained on Google News, and then scraped 36 practice SAT analogies with
answers from various websites. Once I had the analogies, I calculated the
difference between the vectors for each pair of words. For example, for
<code class="language-plaintext highlighter-rouge">King : Man :: Queen : Woman</code>, I would calculate the vector sum for <code class="language-plaintext highlighter-rouge">King −
Man</code> and also for <code class="language-plaintext highlighter-rouge">Queen − Woman</code>. I then computed the cosine distance between
the vectors from the prompt pair of words and the potential answer pairs and
ranked them from lowest to highest. If the model performed well for an
analogy, the correct pair would be the lowest distance from the prompt pair,
otherwise it would be further away.</p>

<p>You can find the Jupyter Notebook used to run the model <a href="/files/sat2vec/Word2Vec%20SAT.ipynb">here</a>
(<a href="https://github.com/agude/agude.github.io/blob/master/files/sat2vec/Word2Vec%20SAT.ipynb">Rendered on Github</a>). You will need the <a href="/files/sat2vec/analogies.json">analogies
data</a> and <a href="/files/sat2vec/vectors.json">pretrained model</a>. The model has been
stripped down to only contain the words that appear in the analogies to save
space.</p>

<h2 id="results">Results</h2>

<p>Here is an example where the model determined the right answer, that is, the correct answer
(in <strong>bold</strong>) is ranked first:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">authenticity : counterfeit</th>
      <th style="text-align: right">Distance</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left"><strong>reliability : erratic</strong></td>
      <td style="text-align: right"><strong>0.758</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">mobility : energetic</td>
      <td style="text-align: right">0.977</td>
    </tr>
    <tr>
      <td style="text-align: left">argument : contradictory</td>
      <td style="text-align: right">0.997</td>
    </tr>
    <tr>
      <td style="text-align: left">reserve : reticent</td>
      <td style="text-align: right">1.009</td>
    </tr>
    <tr>
      <td style="text-align: left">anticipation : solemn</td>
      <td style="text-align: right">1.049</td>
    </tr>
  </tbody>
</table>

<p>Note that <code class="language-plaintext highlighter-rouge">reliability : erratic</code> was the word pair with the lowest
distance, that is, the model predicted that it was the correct answer. Just as
‘counterfeit’ implies lack of authenticity, so ‘erratic’ implies lack of
reliability. The model did in fact succeed in its prediction.</p>

<p>However, the model often failed, as it does for the prompt <code class="language-plaintext highlighter-rouge">paltry :
significance</code>:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">paltry : significance</th>
      <th style="text-align: right">Distance</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">austere : landscape</td>
      <td style="text-align: right">0.803</td>
    </tr>
    <tr>
      <td style="text-align: left">redundant : discussion</td>
      <td style="text-align: right">0.829</td>
    </tr>
    <tr>
      <td style="text-align: left"><strong>banal : originality</strong></td>
      <td style="text-align: right"><strong>0.861</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">oblique : familiarity</td>
      <td style="text-align: right">0.895</td>
    </tr>
    <tr>
      <td style="text-align: left">opulent : wealth</td>
      <td style="text-align: right">0.984</td>
    </tr>
  </tbody>
</table>

<p>Here the correct answer is ranked third. Overall, the model ranked the correct
answer first about 20% of the time. The distribution of answers is as follows:</p>

<p><a href="/files/sat2vec/analogies_ranking.svg"><img src="/files/sat2vec/analogies_ranking.svg" alt="Word2Vec Results on SAT Analogies" /></a></p>

<p>So our model isn’t getting into Berkeley anytime soon; maybe it should try
applying to Stanford instead? (Go Bears!)</p>

<p>The model’s answer for all 36 analogies can be found <a href="/blog/sat2vec/results/">here</a>.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="machine-learning" />
        
      

      

      
      
        <summary type="html"><![CDATA[Could Word2Vec pass the SAT analogies section and get accepted to a good college? I take a pre-trained model and find out!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/sat2vec/Pasternak_The_Night_Before_the_Exam.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/sat2vec/Pasternak_The_Night_Before_the_Exam.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Lab41 Reading Group: Deep Networks with Stochastic Depth</title>
      <link href="https://alexgude.com/blog/lab41-stochastic-depth/" rel="alternate" type="text/html" title="Lab41 Reading Group: Deep Networks with Stochastic Depth" />
      <published>2016-07-11T00:00:00-07:00</published>
      <updated>2016-07-11T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/lab41_stochastic_depth</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/lab41-stochastic-depth/"><![CDATA[<p><a href="https://arxiv.org/abs/1603.09382">Today’s paper</a> is by Gao Huang, Yu Sun, <em>et al.</em><sup style="anchor-name:--fnref-huang" id="fnref:huang"><a href="#fn:huang" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> It introduces
a new way to perturb networks during training in order to improve their
performance. Before I continue, let me first state that this paper is a <strong>real
pleasure to read</strong>; it is concise and extremely well written. It gives an
excellent overview of the motivating problems, previous solutions, and Huang
and Sun’s new approach. I highly recommended giving it a read!</p>

<p>The authors begin by pointing out that deep neural networks have greater
expressive power as compared to shallow networks, that is they can learn more
details and better separate similar classes of objects. For example, a shallow
network might be able to tell cats from dogs, but a deep network has a better
chance of learning to tell a Husky from a Malamute. However, deep networks are
more difficult to train. Huang and Sun list the following issues that appear
when training very deep networks:</p>

<ul>
  <li>
    <p><strong>Vanishing Gradients</strong>: As the gradient information is backpropagated
through the network, it is multiplied by the weights. In a deep network this
multiplication is repeated several times with small weights and so the
information that reaches the earliest layers is often too little to
effectively train the network.</p>
  </li>
  <li>
    <p><strong>Diminishing Feature Reuse</strong>: This is the same problem as the vanishing
gradient, but in the forward direction. Features computed by early layers
are washed out by the time they reach the final layers by the many weight
multiplications in between.</p>
  </li>
  <li>
    <p><strong>Long Training Times</strong>: Deeper networks require a longer time to train than
shallow networks. Training time scales linearly with the size of the network.</p>
  </li>
</ul>

<p>There are many solutions to these problems and the authors propose a new one:
<strong>Stochastic Depth</strong>. In essence what stochastic depth does is randomly bypass
layers in the network while training. They construct their network of
ResBlocks (see image below, and <a href="/blog/lab41-resnet/">my post for more information</a>)
which are a set of convolution layers and a bypass that passes the information
from the previous layer through without any change. With stochastic depth, the
convolution block is sometimes switched off allowing the information to flow
through the layer without being changed, effectively removing the layer from
the network. During testing, all layers are left in and the weights are
modified by their survival probability. This is very similar to how dropout
works, except instead of dropping a single node in a layer the entire layer is
dropped!</p>

<figure>
  
  <a href=" https://alexgude.com/files/resnet/resblock.svg ">
    <img src=" https://alexgude.com/files/resnet/resblock.svg " alt="A diagram of a ResNet block, or ResBlock." decoding="async" />
  </a>
  
  
  
  <figcaption>A ResBlock. The top path is a convolution layer, while the bottom
  path is a pass through. From Huang <em>et al.</em></figcaption>
  
</figure>

<p>Stochastic depth adds a new hyper-parameter, \(p(l)\), the probability of
dropping a layer as a function of its depth. They take \(p(l)\) to be linear
with it equal to 0.0 for the first layer and 0.5 for the last, although other
functions (including a constant) are possible. With this model the expected
depth of a network is effectively reduced by 25% with corresponding reductions
in training time. The authors also show that it reduces the problems
associated with vanishing gradients and diminishing feature reuse, as expected
for a shallower network.</p>

<figure>
  
  <a href=" /files/stochastic-depth//training.png ">
    <img src=" /files/stochastic-depth//training.png " alt="A graph explaining how the network is trained, and the drop
  chance of each layer." decoding="async" />
  </a>
  
  
  
  <figcaption>An example training run on a network with stochastic depth. The red
  and blue bars indicate the probability of dropping a layer, p(l). In this
  example layer 3 and layer 5 have been dropped. From Huang <em>et al.</em></figcaption>
  
</figure>

<p>In addition to aiding in training, the trained networks actually <strong>perform
better</strong> than networks trained without stochastic depth! This is because
stochastic depth, like dropout, acts as a form of regularization, preventing
the network from over training. However, unlike dropout, stochastic depth
works with batch normalization making it a very powerful combination.</p>

<p>The authors demonstrate the new architecture on <a href="https://cave.cs.toronto.edu/kriz/learning-features-2009-TR.pdf">CIFAR-10</a>,
<a href="https://cave.cs.toronto.edu/kriz/learning-features-2009-TR.pdf">CIFAR-100</a>, and the <a href="http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf">Street View House Number dataset</a>
(SVHN). They achieve the lowest published error on CIFAR-10 and CIFAR-100, and
second lowest for SVHN. They also test using a very deep network (1202 layers)
on CIFAR-10 and find that it produces an even better result, the first time a
1000+ layer network has been shown to further reduce the error on CIFAR-10.</p>

<p>The main idea behind stochastic depth is relatively simple, remove some layers
when training to make the network train as if it were shallow, but the results
are surprisingly good. The new networks not only train faster, but they
perform better as well. Further, the idea is compatible with other methods of
improving network training like <a href="https://medium.com/gab41/batch-normalization-what-the-hey-d480039a9e3b">batch normalization</a><sup style="anchor-name:--fnref-ioffe" id="fnref:ioffe"><a href="#fn:ioffe" class="footnote" rel="footnote" role="doc-noteref">2</a></sup>. All in all,
stochastic depth is an essentially free improvement when training a deep
network. I look forward to giving it a shot in my next model!</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:huang">

      <p><span class="citation">Huang, Gao and Sun, Yu and Liu, Zhuang and Sedra, Daniel and Weinberger, Kilian Q. “Deep Networks with Stochastic Depth” <cite>Computer Vision – ECCV 2016</cite>. Edited by Leibe, Bastian and Matas, Jiri and Sebe, Nicu and Welling, Max. Springer International Publishing. 2016. pp. 646–661. doi: <a href="https://doi.org/10.1007/978-3-319-46493-0_39">10.1007/978-3-319-46493-0_39</a>.</span> <a href="#fnref:huang" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:ioffe">

      <p><span class="citation">Ioffe, Sergey and Szegedy, Christian. <a href="https://arxiv.org/abs/1502.03167">“Batch normalization: accelerating deep network training by reducing internal covariate shift”</a> <cite>Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37</cite>. JMLR.org. 2015. pp. 448–456.</span> <a href="#fnref:ioffe" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="reading-group" />
        
          <category term="lab41" />
        
      

      

      
      
        <summary type="html"><![CDATA[Dropout successfully regularizes networks by dropping nodes, but what if we went one step further? Find out how stochastic depth improves your network by dropping whole layers!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/stochastic-depth/header.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/stochastic-depth/header.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
    <entry>
      

      <title type="html">Dragon Farkle: Simulating the End Game</title>
      <link href="https://alexgude.com/blog/simulating-dragon-farkle/" rel="alternate" type="text/html" title="Dragon Farkle: Simulating the End Game" />
      <published>2016-05-30T00:00:00-07:00</published>
      <updated>2016-05-30T00:00:00-07:00</updated>
      <id>https://alexgude.com/blog/simulating_dragon_farkle</id>
      
      
        <content type="html" xml:base="https://alexgude.com/blog/simulating-dragon-farkle/"><![CDATA[<p>Some of my friends came over last week to play board games and while we ate
dinner we played <a href="https://www.zmangames.com/en/products/dragon-farkle/">Dragon Farkle</a> because it was a simple enough
game to not distract us from the meal. The game involves rolling six normal
six-sided dice and a special six-sided die with the following sides: <code class="language-plaintext highlighter-rouge">{0, 0, 0,
0, 1, 2}</code>. On their turn a player rolls the dice to try to recruit soldiers
for their army by scoring points. They score points by rolling various
combinations on the dice (ignoring the special die, which mainly modifies how
many points are awarded) such as three-of-a-kind, straights, and even solitary
1s and 5s. Dice that are part of a scoring combination are removed and the
remaining dice are rerolled to attempt to score more points. If all the dice
are removed, the player gets six more dice and continues. If no scoring
combinations are present in a roll, the roll is a “farkle” and the turn
generally ends.</p>

<p>When a player has enough soldiers (5000 is the minimum specified by the rules)
they may forgo recruiting more to instead attack the dragon. When a player
attacks the dragon they roll exactly as they do when trying to recruit, but
instead of gaining soldiers when they roll scoring combination they lose
soldiers. The special die no longer modifies the number of points awarded but
instead damages the dragon. The player’s turn can end in one of three ways:</p>

<ol>
  <li>
    <p>The player <strong>wins</strong> the game if dragon takes a total of three damage.</p>
  </li>
  <li>
    <p>The player’s turn ends if they roll a farkle <em>and</em> roll a 0 on the special
die. If they roll damage and a farkle they can continue rolling.</p>
  </li>
  <li>
    <p>The player’s turn ends if they run out of soldiers.</p>
  </li>
</ol>

<h2 id="how-many-soldiers">How Many Soldiers?</h2>

<p>The main tension in the game is between spending turns recruiting, and thereby
improving your chances of killing the dragon, and attacking the dragon so that
you can defeat it before your opponents do. Do you attack now, or wait until
you have more soldiers and have a better shot? So of course, the question came
up while playing: “On average, how many soldiers do I need to win?” Although
you can answer this question with pen and paper, I—being an experimental
physicist at heart—decided to simulate the game in order to answer the
question. The notebook that performs the simulation can be found
<a href="/files/dragon_farkle/Dragon%20Farkle%20Simulation.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/dragon_farkle/Dragon%20Farkle%20Simulation.ipynb">rendered on Github</a>). The plotting notebook is
<a href="/files/dragon_farkle/Plot%20Results.ipynb">here</a> (<a href="https://github.com/agude/agude.github.io/blob/master/files/dragon_farkle/Plot%20Results.ipynb">rendered on Github</a>).</p>

<p>Let’s look at how many soldiers you need, on average, to win the game.
To do so, we’ll simulate attacking the dragon and keep track of how many
soldiers are lost before the dragon is defeated. If the player runs out of
soldiers or rolls a farkle we will discard the run. The results are:</p>

<p><a href="/files/dragon_farkle/dragon_farkle_soldier_expectation_value.svg"><img src="/files/dragon_farkle/dragon_farkle_soldier_expectation_value.svg" alt="The average number of soldiers lost during a win." /></a></p>

<p>The various peaks are due to the fact that there are always an integer number
of rolls made in a turn and there is a limited set of scores that can be earned
with each roll. The mean is 1007 soldiers lost before the dragon is defeated,
which is <strong>far</strong> less than the 5000 soldiers that the rules require you to
have before declaring an attack! This suggests that you will almost never run
out of soldiers when attacking the dragon!</p>

<p>Let’s look at exactly how often each of the three end conditions for your turn
are reached. Here we’ll fix the number of soldiers you have, simulate a bunch
of turns, and record the outcome for each. The results are:</p>

<p><a href="/files/dragon_farkle/dragon_farkle_combined_probability.svg"><img src="/files/dragon_farkle/dragon_farkle_combined_probability.svg" alt="The various outcomes of attacking the dragon as a function of the number of
soldiers." /></a></p>

<p>If you go into the turn with 5000 soldiers, you will win about 40% of the time
and you will lose by rolling a farkle about 60% of the time. You will lose by
running out of soldiers less than 1% of the time, and that fraction decreases
exponentially in the number of soldiers!</p>

<p>This is a clear failure of game design; once the minimum number of soldiers is
reached getting more has no effect on the outcome. This removes any possible
strategy where the player might trade time—and hence chances for their
opponents to win—for a higher likelihood of succeeding with their own attacks.
Instead, the only correct strategy is to attack as soon as you can, and keep
doing it as often as you can until you win.</p>]]></content>
      

      
      
      
      
      

      <author>
          <name>Alexander Gude</name>
        
        
      </author>

      
        
          <category term="fun-and-games" />
        
      

      

      
      
        <summary type="html"><![CDATA[How many soldiers do you need to successfully defeat the dragon in Dragon Farkle, and how likely to succeed is your attack? I find out by simulating a game of Dragon Farkle!]]></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://alexgude.com/files/dragon_farkle/st_george_and_the_dragon.jpg" />
        <media:content medium="image" url="https://alexgude.com/files/dragon_farkle/st_george_and_the_dragon.jpg" xmlns:media="http://search.yahoo.com/mrss/" />
      
    </entry>
  
</feed>
