<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-US"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://koulakhilesh.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://koulakhilesh.github.io/" rel="alternate" type="text/html" hreflang="en-US" /><updated>2026-09-27T14:19:20+01:00</updated><id>https://koulakhilesh.github.io/feed.xml</id><title type="html">Akhilesh Koul</title><subtitle>Akhilesh Koul, a data scientist and maker. Projects, writing, and DIY.</subtitle><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><entry><title type="html">How the crowd moves: a typical day on the London Underground</title><link href="https://koulakhilesh.github.io/writing/london-tube-crowding/" rel="alternate" type="text/html" title="How the crowd moves: a typical day on the London Underground" /><published>2026-09-26T00:00:00+01:00</published><updated>2026-09-26T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/london-tube-crowding</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/london-tube-crowding/"><![CDATA[<div class="glyph-hero"><svg class="glyph" viewBox="0 0 44 44" aria-hidden="true"><circle class="sb" cx="22" cy="22" r="17" />
  <circle class="sb" cx="22" cy="22" r="9.5" />
  <circle class="b" cx="22" cy="5" r="2.4" />
  <circle class="b" cx="39" cy="22" r="2.4" />
  <circle class="b" cx="10" cy="34" r="2.4" />
  <circle class="b" cx="28.7" cy="15.3" r="2" />
  <circle class="b" cx="15.3" cy="28.7" r="2" />
  <circle class="r" cx="22" cy="22" r="4.2" /></svg>
</div>

<p>TfL publishes something I didn’t know existed until recently: a “typical day” of crowding for every Underground station, broken into 15-minute slices, for every line that stops there. It isn’t live data, just a model of an ordinary day. It covers the whole network, though, and it comes with a second dataset that estimates how full each train is as it leaves each station.</p>

<p>Most crowding charts I’ve seen answer “where is it busy?” I wanted to know whether this data could show something that <em>moves</em>: whether the rush starts in one place and arrives in another, whether stations have different characters, and what happens inside a train between the platforms. So I pulled the data for 269 stations and looked before deciding what the story was.</p>

<blockquote>
  <p><strong>Data &amp; licence.</strong> Crowding data from the <a href="https://api.tfl.gov.uk">TfL Unified API</a> (<code class="language-plaintext highlighter-rouge">StopPoint/{id}/Crowding/{line}</code>), restricted to the 11 Underground lines and fetched on 26 September 2026. Powered by TfL Open Data. Contains OS data © Crown copyright and database rights 2016 and Geomni UK Map data © and database rights 2019. Used under the <a href="https://tfl.gov.uk/info-for/open-data-users/open-data-policy">TfL Open Data terms</a>. The data is one modelled typical day: no weekday/weekend split, and nothing between 02:00 and 05:00. Distances are straight-line kilometres from Charing Cross. Notebook: <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london_tube_crowding/eda_notebook.ipynb">eda_notebook.ipynb</a>.</p>
</blockquote>

<h2 id="a-typical-day">A typical day</h2>

<p>First, the whole network added together. Every station, every line, 15 minutes at a time.</p>

<figure class="chart">
  <iframe src="/assets/tube/daily-flow.html" title="Total passenger flow across all Tube stations through a typical day" loading="lazy" style="height:460px"></iframe>
  <figcaption>Network flow per 15-minute slice. The gap between 02:00 and 05:00 is missing data, not an empty network.</figcaption>
</figure>

<p>It’s the familiar two-humped camel. The busiest slice of the day is <strong>08:15</strong> (210,474 flow units across the network); the evening peak at <strong>17:45</strong> is about 5% lower (200,336). The evening hump is wider, though, and holds slightly more of the day: <strong>28.1%</strong> of all flow falls between 16:00 and 19:00, against <strong>26.3%</strong> between 07:00 and 10:00. Those six rush hours carry <strong>54.4%</strong> of the flow in a dataset that covers 21 hours.</p>

<p>None of that is surprising. The interesting parts come from pulling this total apart.</p>

<h2 id="where-its-busy-and-a-ranking-i-got-wrong-first">Where it’s busy, and a ranking I got wrong first</h2>

<p>Add up each station across all the lines that serve it and the ranking looks as you’d expect: <strong>Oxford Circus</strong> (316,602), <strong>King’s Cross St Pancras</strong> (284,292), <strong>Green Park</strong>, <strong>Victoria</strong>, <strong>Waterloo</strong>. It’s also very concentrated. The busiest 27 stations, 10% of the network, carry <strong>53.5%</strong> of all the flow.</p>

<p>My first version of this ranking had <strong>Brixton in second place</strong>. Brixton is busy, but not busier than King’s Cross. It turned out that for most stations TfL returns several unlabeled flow series per time slice, in no particular order, and my first parser kept just one of them. Brixton, the end of the Victoria line, happens to have only one series, so it kept everything while the big interchanges lost most of theirs. The same bug made Moorgate look more than twelve times busier in the morning than in the evening. Summing all the series fixed both. The honest caveat is that I still don’t know what each series <em>is</em> (TfL doesn’t label them), so I treat the sum as “typical passenger flow” rather than a confirmed count of entries and exits.</p>

<h2 id="the-rush-reaches-the-centre-last">The rush reaches the centre last</h2>

<p>Here’s the network hour by hour. Press play.</p>

<figure class="chart">
  <iframe src="/assets/tube/hourly-map.html" title="Animated map of station flow across London, hour by hour" loading="lazy" style="height:620px"></iframe>
  <figcaption>Average flow per 15 minutes at each station, one frame per hour from 05:00 to 01:00. Bigger and brighter means busier.</figcaption>
</figure>

<p>It’s pretty, but by volume the centre dominates every frame, so the animation mostly shows the city breathing in and out. To see <em>movement</em> I needed timing rather than size. For each station I found its busiest 15-minute slice between 05:00 and 11:00, then coloured the map by that time.</p>

<figure class="chart">
  <iframe src="/assets/tube/morning-peak-map.html" title="Map of Tube stations coloured by the time of their busiest morning slice" loading="lazy" style="height:620px"></iframe>
  <figcaption>Stations with at least 5,000 flow units a day (203 of them), coloured by when their morning peak hits. Dot size is daily flow.</figcaption>
</figure>

<p>Now the map shows a direction. The earliest stations are all a long way out: Dagenham Heathway, Kingsbury and Queensbury peak at 06:30, Barking and East Ham at 06:45. The centre is dark, peaking around 08:30. Plotted against distance, it’s a clear slope:</p>

<figure class="chart">
  <iframe src="/assets/tube/peak-vs-distance.html" title="Busiest morning slice against distance from Charing Cross" loading="lazy" style="height:460px"></iframe>
  <figcaption>Each dot is a station. Further out, the morning peaks earlier. Hover for names.</figcaption>
</figure>

<table>
  <thead>
    <tr>
      <th>Distance from Charing Cross</th>
      <th>Stations</th>
      <th>Median morning peak</th>
      <th>Median evening peak</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>0–3 km</td>
      <td>47</td>
      <td>08:30</td>
      <td>17:45</td>
    </tr>
    <tr>
      <td>3–8 km</td>
      <td>74</td>
      <td>08:15</td>
      <td>17:30</td>
    </tr>
    <tr>
      <td>8–15 km</td>
      <td>54</td>
      <td>08:00</td>
      <td>17:15</td>
    </tr>
    <tr>
      <td>15 km and beyond</td>
      <td>28</td>
      <td>08:00</td>
      <td>17:00</td>
    </tr>
  </tbody>
</table>

<p>The correlation between morning peak time and distance is <strong>−0.55</strong>. It’s what you’d expect if people leave home at roughly the same time relative to when they need to arrive: the further out you live, the earlier you get on.</p>

<p>I checked one way this could be an artefact. A few central work stations (Temple, Goodge Street, St Paul’s and others) are so quiet in the morning that their “morning peak” is just the last slice of the window, 10:45, which would drag the centre later. Dropping those 7 stations makes the pattern <em>stronger</em>, not weaker: the correlation goes to <strong>−0.63</strong> and the medians in the table don’t move. It isn’t a smooth gradient line by line, either. On the Northern line south of the river, Clapham North peaks at 07:30, earlier than Morden at 08:00. The wave shows up in the aggregate, not neatly at every stop.</p>

<p>The evening column is the one I didn’t expect. I assumed the wave would run in reverse after work, with the centre peaking first and the suburbs later as people got home. Instead the outer stations <em>still</em> peak earlier (17:00 against 17:45, correlation −0.52). I don’t have a clean explanation. Because TfL doesn’t say which flow series is entries and which is exits, I can’t check whether that’s people leaving the suburbs or arriving in them, so I’m leaving it as an open question rather than inventing a story.</p>

<h2 id="home-stations-and-work-stations">Home stations and work stations</h2>

<p>The timing hints at something simpler: some stations are busiest in the morning and others in the evening. Dividing each station’s evening peak by its morning peak splits the network almost cleanly.</p>

<figure class="chart">
  <iframe src="/assets/tube/home-work.html" title="Evening peak divided by morning peak for each station, against distance from the centre" loading="lazy" style="height:480px"></iframe>
  <figcaption>Above the dashed line, the evening is busier; below it, the morning. Log scale. Dot size is daily flow.</figcaption>
</figure>

<p>At the bottom are the home stations. At <strong>Elm Park</strong>, 23 km out on the District line, the evening peak is a tenth of the morning peak. <strong>Pinner</strong> and <strong>Queensbury</strong> are close behind. At the top are the work stations: <strong>Goodge Street</strong>’s evening peak is 10.5 times its morning peak, with <strong>Mansion House</strong>, <strong>Temple</strong> and <strong>Chancery Lane</strong> not far off. Across all 203 stations the correlation between this ratio (on a log scale) and distance is <strong>−0.6</strong>.</p>

<p><em>Try it: <a href="/lab/#tube-day">look up any station’s day in the Lab →</a></em></p>

<p>The exceptions are the fun part. <strong>Canary Wharf</strong>, 7.5 km out, sits among the central stations at 5.1: a second business district doing exactly what the City does. And a handful of stations beyond 20 km are busier in the evening than the morning: <strong>Heathrow</strong> and <strong>Uxbridge</strong>, which are destinations in their own right rather than places people commute from.</p>

<p>That home/work split is also why I think the flow numbers lean towards people <em>entering</em> stations. If the sum counted entries and exits equally, a home station would be busy twice a day, not once. It’s still an inference, though.</p>

<h2 id="night-out-stations-and-a-station-with-no-rush-hour">Night-out stations and a station with no rush hour</h2>

<p>Two more measures pick out different types of station: the share of a station’s day that falls after 20:00, and the share inside the two rush windows.</p>

<p>The median station takes <strong>8.3%</strong> of its daily flow after 20:00. <strong>Covent Garden</strong> takes <strong>32.3%</strong>, <strong>Leicester Square</strong> 30.5%, <strong>Piccadilly Circus</strong> 25.9%. That’s the West End, as you’d guess. At the other end, the median station does <strong>56.3%</strong> of its business in the rush hours, while <strong>Heathrow Terminal 5</strong> does 35.8% and <strong>Terminals 2 &amp; 3</strong> 36.9%. Planes don’t keep office hours.</p>

<figure class="chart">
  <iframe src="/assets/tube/station-types.html" title="Daily flow profiles of four station types, each scaled to its own peak" loading="lazy" style="height:470px"></iframe>
  <figcaption>Each line is scaled to that station's own busiest slice, so the shapes can be compared regardless of size.</figcaption>
</figure>

<p>Four stations, four shapes. Elm Park spikes once in the morning. Goodge Street spikes once in the evening. Leicester Square builds slowly through the day and stays high late. Heathrow is roughly flat, a long plateau with no rush hour at all.</p>

<h2 id="inside-the-train">Inside the train</h2>

<p>Station flow shows where people get on and off. The second dataset shows what happens in between: for each station and direction, a band from 0 to 6 for how full the train is when it leaves. Higher is fuller. That makes it possible to follow a single line and watch a train fill.</p>

<p>Here is the westbound Central line in the morning, from Epping through east London, the City and into the West End.</p>

<figure class="chart">
  <iframe src="/assets/tube/central-line.html" title="Heatmap of train loading on the westbound Central line from Epping to Marble Arch, 06:00 to 10:30" loading="lazy" style="height:640px"></iframe>
  <figcaption>Rows run down the line from Epping to Marble Arch; columns are departure times. Darker means a fuller train.</figcaption>
</figure>

<p>Read the 08:15 column from top to bottom. A train leaving Epping is at band 1. By Leytonstone it’s at 4, by Stratford 5, and leaving <strong>Bethnal Green</strong> and <strong>Liverpool Street</strong> it hits <strong>6</strong>, the top of the scale. Then it empties: 5 leaving Bank, 4 at St Paul’s and Chancery Lane, and by Holborn it’s down to 2, where it stays all the way to Marble Arch. The crowd boards across east London and gets off in the City. On the West End stretch, from Tottenham Court Road to Marble Arch, the train never goes above band 2 all morning.</p>

<p>Across the whole network, <strong>37 of 747</strong> station-to-station stretches ever reach band 6. The ones that stay at band 4 or above the longest:</p>

<table>
  <thead>
    <tr>
      <th>From</th>
      <th>To</th>
      <th>Line</th>
      <th>Hours at band 4+</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Bethnal Green</td>
      <td>Liverpool Street</td>
      <td>Central</td>
      <td>3.5</td>
    </tr>
    <tr>
      <td>King’s Cross St Pancras</td>
      <td>Euston Square</td>
      <td>Hammersmith &amp; City</td>
      <td>3.5</td>
    </tr>
    <tr>
      <td>Moorgate</td>
      <td>Bank</td>
      <td>Northern</td>
      <td>3.5</td>
    </tr>
    <tr>
      <td>Barbican</td>
      <td>Farringdon</td>
      <td>Hammersmith &amp; City</td>
      <td>3.25</td>
    </tr>
    <tr>
      <td>Bermondsey</td>
      <td>London Bridge</td>
      <td>Jubilee</td>
      <td>3.25</td>
    </tr>
    <tr>
      <td>Southwark</td>
      <td>Waterloo</td>
      <td>Jubilee</td>
      <td>3.25</td>
    </tr>
  </tbody>
</table>

<p>Does train fullness show the same inward wave as station flow? Partly. For each inbound stretch I took a weighted average of <em>when</em> in the morning it’s full. That time does get later nearer the centre (correlation <strong>−0.41</strong>): about 08:00 for stretches 8–15 km out, 08:09 at 3–8 km, 08:22 in the centre. But the outermost stretches, beyond 15 km, come in at 08:08, <em>later</em> than the band inside them. A 0–6 scale is coarse, and trains near the ends of lines are fairly empty whenever you measure them, so I wouldn’t read much into that last row. The wave is clearest in the station data.</p>

<h2 id="what-the-data-says">What the data says</h2>

<ul>
  <li><strong>The morning rush moves inward.</strong> Across the network, outer stations peak around 08:00 and central ones around 08:30. The pattern holds in aggregate even though individual lines are bumpier.</li>
  <li><strong>Stations have characters.</strong> Home stations spike in the morning, work stations in the evening, and the ratio between them tracks distance from the centre, with Canary Wharf, Heathrow and Uxbridge as the exceptions you’d hope a good measure would catch.</li>
  <li><strong>The West End and the airport break the pattern.</strong> Covent Garden does a third of its business after 20:00; Heathrow barely has a rush hour.</li>
  <li><strong>You can watch a train fill up.</strong> The morning Central line goes from band 1 at Epping to the top band at Bethnal Green and back down to band 2 by Holborn.</li>
  <li><strong>The evening doesn’t run the wave in reverse.</strong> Outer stations still peak first after work, and without labelled entries and exits I can’t say why.</li>
</ul>

<p>The part I’ll remember is the bug. My first chart said Brixton was the second-busiest station on the Underground, and it looked plausible enough that I nearly kept it. The station-by-station timing held up once I summed the series properly, but the ranking only became believable after I stopped trusting the first number. The <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london_tube_crowding/eda_notebook.ipynb">notebook is here</a> if you want to follow your own line.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="London" /><category term="Open Data" /><category term="Plotly" /><category term="Transport" /><summary type="html"><![CDATA[TfL publishes a 'typical day' of passenger flow for every Tube station, in 15-minute slices. I pulled it for 269 stations and asked whether you can watch the crowd move. You can: the morning rush reaches the centre last, stations split cleanly into home and work, and on the Central line you can see a train fill up across east London and empty out in the City.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/london-tube-crowding.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/london-tube-crowding.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Building a knowledge graph of the Bhagavad Gita</title><link href="https://koulakhilesh.github.io/writing/building-a-knowledge-graph-of-the-gita/" rel="alternate" type="text/html" title="Building a knowledge graph of the Bhagavad Gita" /><published>2026-08-28T00:00:00+01:00</published><updated>2026-08-28T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/building-a-knowledge-graph-of-the-gita</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/building-a-knowledge-graph-of-the-gita/"><![CDATA[<div class="glyph-hero"><svg class="glyph" viewBox="0 0 44 44" aria-hidden="true"><polyline class="k" points="12,29 22,13 33,27 12,29" />
  <polyline class="k" points="33,27 27,39" />
  <rect class="r" x="7" y="24" width="10" height="10" />
  <polygon class="y" points="33,19 40,32 26,32" />
  <circle class="b" cx="22" cy="13" r="5" />
  <circle class="b" cx="27" cy="39" r="3" /></svg>
</div>

<p><a href="/writing/the-shape-of-the-gita/">The shape of the Gita</a> showed the pictures: a map of 700 verses, the regions the text falls into, the way its three voices each use different words. This post is the plumbing underneath those pictures. Before you can ask a text a statistical question you have to turn it into something a machine can hold, and I wanted that step to be honest: nothing appearing from nowhere, and the whole thing rebuildable from the source with one command.</p>

<p>What comes out is a graph of <strong>5,314 nodes</strong> and <strong>21,947 relationships</strong>, and every single one of them can be traced back to a verse file.</p>

<blockquote>
  <p><strong>Method note.</strong> The source is my digital edition of the Gita: 701 verse records, each a markdown file with the chapter and verse number, the Sanskrit, a transliteration, a word-by-word gloss, and an English translation. Everything below is deterministic and idempotent (re-running rebuilds the same graph with no duplicates), and the pure parsing and graph-building logic is covered by 89 unit tests that touch neither the database nor the network. Code and the full ontology live in the <a href="https://github.com/koulakhilesh/CodePlayground/tree/main/gita-knowledge-graph">gita-knowledge-graph folder</a> on GitHub.</p>
</blockquote>

<h2 id="the-verse-is-the-unit">The verse is the unit</h2>

<p>Everything hangs off the verse. Parsing one is deliberately boring, exact string and regular-expression work rather than anything clever: pull the chapter and verse number from the front matter, take the text under the <code class="language-plaintext highlighter-rouge">## Translation</code> heading, and read the word-by-word table under <code class="language-plaintext highlighter-rouge">## Word Meanings</code>. That gives the spine of the graph: one <strong>Text</strong>, its <strong>18 chapters</strong>, and <strong>701 verses</strong>, chained in reading order so you can walk the book verse by verse.</p>

<p>Two facts about each verse are worth a little care. The first is who speaks it. The Gita marks a change of speaker with a prefix, <em>“The Blessed Lord said,”</em> or <em>“Arjuna said,”</em>, but most verses carry no prefix at all because the speaker has not changed. So the parser reads the prefix when it is there and otherwise inherits the previous speaker, which is exactly how a human reads it. That resolves all 701 verses to one of <strong>four voices</strong> and, from the speaker, an addressee.</p>

<p>The second is the honorifics. The Gita almost never uses plain names; it calls Arjuna <em>Partha</em> and <em>Dhananjaya</em>, and Krishna <em>Hrishikesha</em> and <em>Madhava</em>. I tag those with a rule-based matcher seeded from a fixed list of <strong>22 epithets</strong>, not a statistical name-finder. That is a deliberate choice: the list is small and known, so a rule is both more accurate and more honest than a model that might hallucinate a name.</p>

<h2 id="two-vocabularies">Two vocabularies</h2>

<p>A verse has two texts, and I did not want to throw either away. So the graph carries two parallel vocabularies.</p>

<p>The <strong>English layer</strong> comes from the translation. A standard NLP pipeline (spaCy) lemmatises it and keeps the nouns and verbs, which become <strong>1,171 distinct terms</strong> joined to their verses by <strong>6,156</strong> weighted links. This is the statistical layer, the one that later feeds themes.</p>

<p>The <strong>Sanskrit layer</strong> comes from the word-by-word gloss. This one needed more care, because a word-by-word table is full of grammatical noise: pronouns, particles, and speaker markers that carry no meaning. So the parser drops those with a stoplist and normalises inflected surface forms back toward their root, so that <em>karmani</em>, <em>karmana</em>, and <em>karma-phala</em> all land on <em>karma</em>. That leaves <strong>3,340 Sanskrit terms</strong> linked to their verses <strong>7,995 times</strong>. The design rule I held to: English drives the embeddings later, because it is reproducible, but Sanskrit drives <em>meaning</em>, because it is the source language.</p>

<h2 id="from-words-to-ideas">From words to ideas</h2>

<p>Two vocabularies are still just words. The next layer turns them into ideas, and it does so deterministically, with no model guessing.</p>

<p><strong>Themes</strong> are 13 broad topics (karma, dharma, bhakti, jnana, yoga, and so on), each defined as a set of English lemmas. A verse links to a theme when it uses those lemmas, and the strength of the link is just how often. <strong>Concepts</strong> are 22 philosophical categories grounded the other way, in the Sanskrit terms: a verse expresses <em>atman</em> when it contains Sanskrit words rooted in <em>atman</em>. The two layers meet on the nine names they share (karma, dharma, yoga, bhakti, jnana, moksha, atman, brahman, guna), bridged so a query can cross from the English index into the Sanskrit ontology.</p>

<p>Because every verse now carries themes, you can ask which ideas keep company. This is node similarity over the shared verses: two themes score highly when they tend to appear in the same verses.</p>

<figure class="chart">
  <iframe src="/assets/gita/analysis_theme_correlation.html" title="Heatmap of which Gita themes co-occur, measured by Jaccard similarity over shared verses" loading="lazy" style="height:620px"></iframe>
  <figcaption>Which themes travel together, by how many verses they share. Brighter means more overlap.</figcaption>
</figure>

<h2 id="the-cast-the-conches-and-the-glories">The cast, the conches, and the glories</h2>

<p>The nicest part to build was the narrative layer, because it is pulled straight out of the glosses rather than curated by hand. Scan the word meanings for known names and you recover the <strong>cast</strong>: 18 characters, from Arjuna and Krishna down to minor warriors, joined to the verses that name them by <strong>310</strong> links.</p>

<figure class="chart">
  <iframe src="/assets/gita/analysis_character_network.html" title="Co-occurrence network of characters named in the Gita" loading="lazy" style="height:640px"></iframe>
  <figcaption>Two characters are linked when a verse names them together. Node size is centrality in that network.</figcaption>
</figure>

<p>Two smaller entity layers come from the same glosses. Chapter 1 names six war-conches, and the graph knows which warrior blows each one:</p>

<table>
  <thead>
    <tr>
      <th>Conch</th>
      <th>Sounded by</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Panchajanya</td>
      <td>Krishna</td>
    </tr>
    <tr>
      <td>Devadatta</td>
      <td>Arjuna</td>
    </tr>
    <tr>
      <td>Paundra</td>
      <td>Bhima</td>
    </tr>
    <tr>
      <td>Anantavijaya</td>
      <td>Yudhishthira</td>
    </tr>
    <tr>
      <td>Sughosha</td>
      <td>Nakula</td>
    </tr>
    <tr>
      <td>Manipushpa</td>
      <td>Sahadeva</td>
    </tr>
  </tbody>
</table>

<p>And Chapter 10, where Krishna lists his divine glories, is captured as the 13 verses that explicitly declare <em>“I am…”</em> (in Sanskrit, <em>asmi</em>). I kept that layer strict on purpose: it records <em>that</em> Krishna declares a glory, but it does not try to pair each “I am” with its object, because the Sanskrit word order there is inconsistent and any automatic pairing would be guessing. I would rather record less and have all of it be right.</p>

<h2 id="meaning-by-number">Meaning by number</h2>

<p>The last layer is the one the map is built on. Every verse gets a 768-dimensional embedding of its <strong>English translation only</strong>, from a pinned <code class="language-plaintext highlighter-rouge">all-mpnet-base-v2</code> model, and I compare every pair by cosine similarity. Rather than keep every faint resemblance, I calibrated a threshold: the notebook tries several cutoffs, checks how much of the text stays connected and how many cross-chapter theme links survive, and settles on 0.55, which keeps <strong>2,260</strong> verse-to-verse similarity links.</p>

<p>One deliberate piece of friction here: those similarity edges are the only part of the build that is not written automatically. The notebook stops and makes me eyeball a sample of the pairs and approve them before it loads them. It is the one place where the graph makes a judgement about meaning, so it is the one place I kept a human in the loop.</p>

<h2 id="the-whole-thing-and-why-it-is-built-this-way">The whole thing, and why it is built this way</h2>

<p>Stack all of that up and you get the graph. Here is its backbone, everything except the two dense word layers, which would otherwise drown the picture:</p>

<figure class="chart">
  <iframe src="/assets/gita/gita_graph_backbone.html" title="The structural backbone of the Gita knowledge graph" loading="lazy" style="height:680px"></iframe>
  <figcaption>Text, chapters, verses, speakers, themes, concepts, characters, conches, and the similarity web between verses. The word layers are hidden here.</figcaption>
</figure>

<p>The one principle that kept the whole thing honest is that every node and edge has a <strong>provenance</strong>, and it is always one of four kinds. It is a <em>seed</em> (a hand-curated constant, like the chapter names or the theme definitions), something <em>extracted</em> directly from a verse (the speaker, the epithets, the Sanskrit terms, the conches), something <em>derived</em> deterministically from what was extracted (the themes and concepts), or something <em>computed</em> by the pinned model (the embeddings and the similarity). Nothing in the graph is a guess dressed as a fact, and if you disagree with a link you can always trace it back to the verse it came from.</p>

<p>That discipline is the actual point. The pictures in <a href="/writing/the-shape-of-the-gita/">the shape of the Gita</a> are only worth looking at because the thing underneath them is dull, checkable, and rebuilds from the text with one command. The analysis was the fun half; this is the half that makes it trustworthy.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="NLP" /><category term="Knowledge Graph" /><category term="Neo4j" /><category term="spaCy" /><summary type="html"><![CDATA[The companion to the map: how 700 markdown verse files became a graph of 5,314 nodes and nearly 22,000 relationships, deterministically and rebuildably, with the English translation and the Sanskrit word meanings both wired in.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/building-a-knowledge-graph-of-the-gita.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/building-a-knowledge-graph-of-the-gita.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The shape of the Gita: mapping 700 verses with a knowledge graph</title><link href="https://koulakhilesh.github.io/writing/the-shape-of-the-gita/" rel="alternate" type="text/html" title="The shape of the Gita: mapping 700 verses with a knowledge graph" /><published>2026-08-28T00:00:00+01:00</published><updated>2026-08-28T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/the-shape-of-the-gita</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/the-shape-of-the-gita/"><![CDATA[<div class="glyph-hero"><svg class="glyph" viewBox="0 0 44 44" aria-hidden="true"><circle class="b" cx="13" cy="14" r="2" />
  <circle class="b" cx="17" cy="11" r="2" />
  <circle class="b" cx="10" cy="18" r="2" />
  <circle class="b" cx="16" cy="17" r="2" />
  <circle class="r" cx="31" cy="18" r="2" />
  <circle class="r" cx="34" cy="13" r="2" />
  <circle class="r" cx="29" cy="12" r="2" />
  <circle class="r" cx="34" cy="22" r="2" />
  <circle class="y" cx="20" cy="31" r="2" />
  <circle class="y" cx="25" cy="34" r="2" />
  <circle class="y" cx="17" cy="34" r="2" />
  <circle class="y" cx="24" cy="29" r="2" /></svg>
</div>

<p>The Bhagavad Gita is a conversation of about 700 verses, 18 chapters, one battlefield. I have been building a verse-by-verse digital edition of it as <a href="/projects/the-gita-project/">The Gita Project</a>, and once every verse had its Sanskrit, a transliteration, a word-by-word gloss, and an English translation, a different question started nagging at me. Not <em>what does it say</em>, which people have argued about for two thousand years, but a smaller and more answerable one: <strong>what does the text look like when you map it?</strong></p>

<p>So I loaded all of it into a graph database, gave every verse a numeric fingerprint, and let a few standard algorithms draw the picture. This post is what came back. None of it settles a single theological argument. But the data does say some concrete things, and one of them, about how the three speakers talk, surprised me.</p>

<blockquote>
  <p><strong>Method note.</strong> The source is my own digital edition of the Gita: 701 verse records (the text is traditionally counted as 700; this edition also carries Arjuna’s opening question in Chapter 13, which some recensions omit, so it runs to 701), each with its English translation and its Sanskrit Word Meanings. I loaded them into a local <strong>Neo4j</strong> graph, then layered themes, Sanskrit-grounded concepts, the character cast, and semantic similarity on top. Every verse also gets a 768-dimensional embedding of its <strong>English translation only</strong>, from a pinned <code class="language-plaintext highlighter-rouge">sentence-transformers/all-mpnet-base-v2</code> model, and verse-to-verse similarity edges are kept above a calibrated cosine threshold of 0.55 (2,260 pairs). The graph analytics use the <strong>Neo4j Graph Data Science</strong> library. The whole thing is deterministic and rebuildable; code and the full ontology live in the <a href="https://github.com/koulakhilesh/CodePlayground/tree/main/gita-knowledge-graph">gita-knowledge-graph folder</a> on GitHub. How the graph itself is built is a companion post, <a href="/writing/building-a-knowledge-graph-of-the-gita/">Building a knowledge graph of the Bhagavad Gita</a>. Numbers in this post come straight from those notebooks.</p>
</blockquote>

<h2 id="a-map-of-700-verses">A map of 700 verses</h2>

<p>Every verse is a point in 768-dimensional space. That is impossible to look at, so I squashed it down to two dimensions with t-SNE, where verses that mean similar things end up near each other. Then I coloured each point by the community a graph algorithm (Louvain) finds in the similarity network, and labelled each region by the theme that is most distinctive to it.</p>

<figure class="chart">
  <iframe src="/assets/gita/map_verse_annotated.html" title="A t-SNE map of all 701 Gita verses, coloured by semantic community and labelled by each region&#39;s most distinctive theme" loading="lazy" style="height:700px"></iframe>
  <figcaption>Every dot is a verse. Nearby dots mean similar things. Colours are the semantic communities; labels name each region's most distinctive theme. Hover a point for its verse number and text.</figcaption>
</figure>

<p>The algorithm splits the text into <strong>47 communities</strong>, and it is a genuinely tight partition: <strong>77%</strong> of all similarity links fall inside a community rather than between communities, and the modularity score is <strong>0.69</strong>. A second, independent algorithm (Leiden) recovers a closely matching split of 49 groups, so this is not an artefact of one method.</p>

<p>The communities also do not respect the chapter numbers, which is the part I keep returning to. Colour the exact same map by chapter instead, and the colours smear across the whole plane rather than forming 18 neat blocks:</p>

<figure class="chart">
  <iframe src="/assets/gita/map_verse_chapters.html" title="The same verse map coloured by chapter number, showing that chapters do not form tidy clusters" loading="lazy" style="height:700px"></iframe>
  <figcaption>The same 701 verses, now coloured 1 to 18 by chapter. If chapters were self-contained topics, you would see 18 blocks. You do not.</figcaption>
</figure>

<p>The largest semantic communities each span ten to sixteen different chapters. The Gita returns to its core ideas again and again, in different chapters, in language similar enough that a machine groups them without ever being told what a chapter is. The chapter divisions organise the reading; the topics ignore them.</p>

<h2 id="what-each-region-is-about">What each region is about</h2>

<p>Naming a cluster is harder than finding it, because the obvious method fails. If you just ask which theme is heaviest in each community, almost every region comes back “karma”, because action is discussed nearly everywhere. So instead I measured <strong>lift</strong>: how much more a theme appears in a community than in the text overall. That surfaces what makes a region distinctive rather than what is simply common.</p>

<figure class="chart">
  <iframe src="/assets/gita/map_community_theme_signature.html" title="Heatmap of theme share for each of the largest verse communities" loading="lazy" style="height:560px"></iframe>
  <figcaption>The theme signature of each of the eight largest communities. Read a column top to bottom to see what that region dwells on.</figcaption>
</figure>

<p>Pick the verse nearest each community’s centre and the regions come into focus. The biggest community, 90 verses, is a <strong>devotion</strong> cluster; its centre is 12.6, <em>“But to those who worship Me, renouncing all actions in Me…”</em>. A second region of 73 verses is pure <strong>karma-yoga</strong>; its centre is 3.30, <em>“Renouncing all actions in Me, with the mind centered on the Self…”</em>. A third, 44 verses, is about the <strong>senses and the mind</strong>, centred on 2.55, <em>“When a man completely casts off, O Arjuna, all the desires of the mind…”</em>.</p>

<p>One honest caveat before anyone reads too much into it: these are <em>semantic</em> clusters, grouped by how the English translations read, not a doctrinal map. They line up loosely with the tradition’s sense that the Gita moves through action, knowledge, and devotion, but I did not build them to prove that, and I would not lean on them to.</p>

<h2 id="the-centre-of-gravity">The centre of gravity</h2>

<p>If verses are a network, some sit closer to the middle than others. PageRank, the algorithm that made Google, scores a verse highly when many other well-connected verses resemble it. Run it on the similarity network and the single most central verse in the whole Gita is <strong>12.6</strong>, followed by <strong>12.7</strong>, both from Chapter 12, the <em>Bhakti Yoga</em> chapter on devotion.</p>

<figure class="chart">
  <iframe src="/assets/gita/analysis_pagerank_top_verses.html" title="The fifteen verses with the highest PageRank in the Gita similarity network" loading="lazy" style="height:560px"></iframe>
  <figcaption>The verses most central to the semantic web. Chapter 12's devotional summations sit at the top.</figcaption>
</figure>

<p>I want to be careful about what this means. PageRank rewards a verse for echoing many others, so the “centre” is really the text’s most <em>representative</em> language, the lines that restate its recurring promise most plainly. That does not make them the most important verses. It only means that if you had to pick the ones the rest of the book most sounds like, the machine points at Chapter 12’s devotional summations.</p>

<h2 id="three-voices-three-vocabularies">Three voices, three vocabularies</h2>

<p>The Gita is a dialogue, and the graph knows who speaks each verse. Four voices carry it, but the floor is not shared evenly at all.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value">82<span class="stat-unit">%</span></div>
    <div class="stat-label">spoken by Krishna (574 verses)</div>
  </div>
  <div class="stat">
    <div class="stat-value">12<span class="stat-unit">%</span></div>
    <div class="stat-label">Arjuna (86 verses)</div>
  </div>
  <div class="stat">
    <div class="stat-value">6<span class="stat-unit">%</span></div>
    <div class="stat-label">Sanjaya, the narrator (40 verses)</div>
  </div>
</div>

<figure class="chart">
  <iframe src="/assets/gita/speaker_share.html" title="Verses spoken by each voice in the Gita" loading="lazy" style="height:340px"></iframe>
</figure>

<p>The interesting question is not who talks most, it is whether they talk <em>differently</em>. To test that I used <strong>keyness</strong>, the standard corpus-linguistics measure: for each speaker, a log-likelihood test flags the words they use far more often than the other voices do. It is the same idea behind “signature words”. The result splits the three main speakers cleanly, and it reads like the story itself.</p>

<figure class="chart">
  <iframe src="/assets/gita/speaker_keyness.html" title="Signature words for each speaker, ranked by Dunning log-likelihood keyness" loading="lazy" style="height:560px"></iframe>
  <figcaption>Content words each voice over-uses relative to the others, by log-likelihood.</figcaption>
</figure>

<p><strong>Krishna</strong> speaks the language of the teacher: his signature words are <em>action, self, sacrifice, knowledge, attachment, intellect</em>. <strong>Arjuna</strong>, the warrior having a breakdown at the edge of a war he does not want to fight, over-uses <em>kill, family, battle, destruction</em>, and, tellingly, <em>mouth</em> and <em>tooth</em>, the vocabulary of his terrifying vision of the divine in Chapter 11. <strong>Sanjaya</strong>, the narrator relaying the scene to a blind king, deals in <em>son, conch, archer, army, king</em>, the furniture of the battlefield he is describing. Nobody labelled these voices by role; a frequency test pulled the teacher, the panicking student, and the war reporter apart on its own.</p>

<p>There is a quieter signal in the same data. Arjuna leans into the theme of <em>dharma</em>, duty, nearly three times as heavily as the text does on average (a lift of about 2.9), which is exactly the knot he is tied in. The machine found his crisis by counting words.</p>

<h2 id="the-honest-limits">The honest limits</h2>

<p>A few things this cannot do, stated plainly so the pictures above are not oversold. The embeddings are built from the <strong>English translation only</strong>, so the map reflects one translator’s choices as much as the Sanskrit; a different translation would move the dots. The keyness test runs over content words (nouns and verbs), so it captures <em>what</em> each voice talks about better than <em>how</em> they phrase it. The community detection is semantic, not doctrinal, and the second algorithm agrees with the first on only about 60% of the exact assignments (adjusted Rand index around 0.58), so treat the fine boundaries as soft. And a t-SNE map distorts distance to fit everything on a page; read it for neighbourhoods, not for precise gaps.</p>

<p>What survives all of that is small but real. The Gita’s topics run across its chapters rather than staying inside them, its most representative language is devotional, and its three voices use measurably different words that match the roles they play. None of it needed me to interpret a single verse.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="NLP" /><category term="Knowledge Graph" /><category term="Neo4j" /><category term="Plotly" /><summary type="html"><![CDATA[I turned all 700 verses of the Bhagavad Gita into a knowledge graph and a set of embeddings, then asked what the text looks like when you map it: where its regions are, which verses sit at the centre, and how its three voices each speak a measurably different language.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/the-shape-of-the-gita.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/the-shape-of-the-gita.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The words the Gita repeats: counting its Sanskrit vocabulary</title><link href="https://koulakhilesh.github.io/writing/the-words-the-gita-repeats/" rel="alternate" type="text/html" title="The words the Gita repeats: counting its Sanskrit vocabulary" /><published>2026-08-28T00:00:00+01:00</published><updated>2026-08-28T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/the-words-the-gita-repeats</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/the-words-the-gita-repeats/"><![CDATA[<div class="glyph-hero"><svg class="glyph" viewBox="0 0 44 44" aria-hidden="true"><rect class="r" x="4" y="12" width="5" height="26" />
  <rect class="b" x="12" y="21" width="5" height="17" />
  <rect class="b" x="20" y="27" width="5" height="11" />
  <rect class="b" x="28" y="31" width="5" height="7" />
  <rect class="b" x="36" y="34" width="5" height="4" /></svg>
</div>

<p>The <a href="/writing/the-shape-of-the-gita/">map</a> and the <a href="/writing/building-a-knowledge-graph-of-the-gita/">build</a> both leaned on the English translation, because that is what the embeddings are made from. But the graph also carries the other half of every verse: the word-by-word Sanskrit gloss, normalized into <strong>3,340 Sanskrit terms</strong>. This post does the least glamorous thing you can do to a vocabulary, which is count it, and finds that the counting says more than I expected.</p>

<blockquote>
  <p><strong>Method note.</strong> The Sanskrit terms come from the per-verse Word Meanings tables, with particles and pronouns dropped and inflected forms normalized toward their roots. The graph records that a term appears in a verse, but not how many times, so throughout this post <strong>frequency means document frequency</strong>: how many of the 701 verses a term shows up in. A “hapax” is therefore a term that appears in exactly one verse. Two honest caveats up front: the normalization is imperfect (you will see a stray inflected form or pronoun below), and document-frequency counting slightly inflates how rich the vocabulary looks. Code lives in the <a href="https://github.com/koulakhilesh/CodePlayground/tree/main/gita-knowledge-graph">gita-knowledge-graph folder</a> on GitHub; numbers come straight from the notebook.</p>
</blockquote>

<h2 id="the-most-repeated-words-are-the-ideas">The most repeated words are the ideas</h2>

<p>Rank the Sanskrit terms by how many verses contain them, and the top of the list is not incidental vocabulary. It is the philosophy itself.</p>

<table>
  <thead>
    <tr>
      <th>Sanskrit term</th>
      <th>Gloss</th>
      <th style="text-align: right">Verses</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>karma</td>
      <td>action</td>
      <td style="text-align: right">92</td>
    </tr>
    <tr>
      <td>jnana</td>
      <td>knowledge</td>
      <td style="text-align: right">60</td>
    </tr>
    <tr>
      <td>atman</td>
      <td>the self</td>
      <td style="text-align: right">55</td>
    </tr>
    <tr>
      <td>yoga</td>
      <td>discipline</td>
      <td style="text-align: right">48</td>
    </tr>
    <tr>
      <td>manas</td>
      <td>mind</td>
      <td style="text-align: right">42</td>
    </tr>
    <tr>
      <td>brahma</td>
      <td>the absolute</td>
      <td style="text-align: right">38</td>
    </tr>
    <tr>
      <td>indriya</td>
      <td>the senses</td>
      <td style="text-align: right">26</td>
    </tr>
    <tr>
      <td>kama</td>
      <td>desire</td>
      <td style="text-align: right">23</td>
    </tr>
    <tr>
      <td>tapah</td>
      <td>austerity</td>
      <td style="text-align: right">22</td>
    </tr>
  </tbody>
</table>

<p><em>Karma</em> appears in <strong>92</strong> of the 701 verses, more than one verse in eight. Under it sit <em>jnana</em>, <em>atman</em>, <em>yoga</em>, <em>manas</em>, <em>brahman</em>: exactly the concepts the tradition says the Gita is about. You could have guessed the list. What I did not expect was that a blind word-count would recover it with no idea of meaning at all.</p>

<p>I am keeping this honest, so: the raw list also has a couple of passengers. The pronoun <em>mayi</em> (“unto me”) slips past the stoplist, the name <em>Arjuna</em> is frequent for obvious reasons, and you can spot <em>buddhih</em> sitting separately from <em>buddhi</em>, an inflected form the normaliser missed. None of that changes the shape of the top of the list, but it is the kind of thing that would, if I swept it under the rug, quietly make the numbers a lie.</p>

<h2 id="it-behaves-like-a-language">It behaves like a language</h2>

<p>Word frequencies in natural language follow Zipf’s law: the <em>n</em>-th most common word appears about proportionally to 1/<em>n</em>, so a log-log plot of frequency against rank falls on roughly a straight line. Does a 700-verse Sanskrit text, counted by document frequency, do the same?</p>

<figure class="chart">
  <iframe src="/assets/gita/sanskrit_zipf.html" title="Log-log plot of Sanskrit term frequency against frequency rank, with a fitted line" loading="lazy" style="height:520px"></iframe>
  <figcaption>Every dot is a Sanskrit term. The straight-ish line in log-log space is Zipf's law showing up. Hover for the word and its gloss.</figcaption>
</figure>

<p>It does, with a fitted slope of about <strong>-0.69</strong>. That is shallower than the textbook -1, which is expected: document frequency flattens the very top (a word can only be counted once per verse, no matter how often it is chanted inside that verse), and the text is short. But the law is unmistakably there. The Gita’s Sanskrit is not a special code; it distributes its words the way language does.</p>

<h2 id="most-words-appear-once">Most words appear once</h2>

<p>The flip side of a Zipf curve is a long, thin tail, and here it is dramatic. <strong>56%</strong> of all Sanskrit terms, 1,865 of the 3,340, appear in exactly one verse. Three quarters appear in two verses or fewer. The vocabulary is a small hard core of repeated concepts sitting on top of a huge scatter of words used once and never again.</p>

<p>You can watch that happen by reading the text in order and counting how many <em>new</em> words each verse brings in:</p>

<figure class="chart">
  <iframe src="/assets/gita/sanskrit_growth.html" title="Cumulative count of distinct Sanskrit terms as the text is read in order" loading="lazy" style="height:460px"></iframe>
  <figcaption>Distinct Sanskrit terms seen so far, verse by verse. The curve barely bends: the Gita keeps minting new words to the end.</figcaption>
</figure>

<p>The curve hardly flattens. By the halfway point of the book only about <strong>1,970</strong> of the 3,340 terms have appeared, so the second half is still introducing more than a thousand new ones. In quantitative-linguistics terms the vocabulary-growth (Heaps’) exponent is about <strong>0.88</strong>, high for a text this size.</p>

<p>Here is where I have to be careful, because it is tempting to read that as “the Gita is extraordinarily rich” and stop. Part of the richness is real: it ranges over metaphysics, ethics, devotion, and a battlefield. But part of it is Sanskrit’s morphology, a single root throws off many inflected surface forms, and my normalization only pulls them <em>toward</em> their roots, not perfectly onto them (remember <em>buddhih</em> and <em>buddhi</em>). So the type count is inflated, and the honest reading is: the Gita has a genuinely broad vocabulary, made to look even broader by the language’s grammar and my imperfect tidying of it.</p>

<h2 id="which-chapters-are-densest">Which chapters are densest</h2>

<p>Finally, not every chapter carries the same weight of vocabulary per verse.</p>

<figure class="chart">
  <iframe src="/assets/gita/sanskrit_richness_by_chapter.html" title="Sanskrit terms per verse for each of the 18 chapters" loading="lazy" style="height:460px"></iframe>
  <figcaption>Distinct Sanskrit terms per verse, by chapter. Colour is the chapter's total distinct-term count.</figcaption>
</figure>

<p>The densest chapter is <strong>16</strong> (<em>Daivasura Sampada Vibhaga</em>, the divine and demoniac natures) at about <strong>17 terms per verse</strong>, a chapter that works by piling up lists of qualities. Chapter <strong>1</strong>, Arjuna’s despair on the field, is close behind at nearly 14, thick with the names of warriors and weapons rather than concepts. The leanest is Chapter <strong>7</strong> at about 9, where Krishna settles into steady, repetitive teaching and reuses the same core words. That is enumeration, not depth: chapter 16 lists where chapter 7 repeats.</p>

<p>None of this interprets a single verse; it only counts them. But the count lands on something real: the words the Gita says most often are its core ideas, the frequencies follow the same law as any language, and it keeps reaching for new words right to the last chapter.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="NLP" /><category term="Linguistics" /><category term="Sanskrit" /><category term="Plotly" /><summary type="html"><![CDATA[The first two posts leaned on the English translation. This one counts the Sanskrit: 3,340 words from the word-by-word glosses, where they follow the same frequency laws as any language, and where the count quietly tells you what the text is actually about.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/the-words-the-gita-repeats.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/the-words-the-gita-repeats.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">A record-hot summer arrives: how London’s reservoirs are taking it</title><link href="https://koulakhilesh.github.io/writing/london-reservoirs-hot-summer/" rel="alternate" type="text/html" title="A record-hot summer arrives: how London’s reservoirs are taking it" /><published>2026-08-15T00:00:00+01:00</published><updated>2026-08-15T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/london-reservoirs-hot-summer</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/london-reservoirs-hot-summer/"><![CDATA[<p><a href="/writing/london-reservoirs/">The first reservoir post</a> stopped on <strong>31 May 2026</strong>, with both of London’s big storage groups brimming: the Lower Lee at 95% of capacity, the Lower Thames at 90%. A comfortable place to end a story about water.</p>

<p>Then the summer came. The Met Office now says a <a href="https://www.metoffice.gov.uk/blog/2026/warmest-uk-summer-on-record-increasingly-likely-as-temperatures-stay-well-above-average">record-warm UK summer is increasingly likely</a>, with temperatures staying well above average. So I did the obvious thing: grabbed the June and July readings, added them to the same 37-year daily series, and looked. I wanted to resist the tidy headline (“record heat drains the reservoirs!”) and let the readings talk first, because they usually say something more interesting than the slogan.</p>

<blockquote>
  <p><strong>Data &amp; licence.</strong> London reservoir levels via the <a href="https://data.london.gov.uk/dataset/london-reservoir-levels-24ry5">London Datastore</a>, sourced from the Environment Agency’s <a href="https://www.gov.uk/government/collections/water-situation-reports-for-england">water situation reports</a>. Public sector information licensed under the <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/2/">Open Government Licence v2.0</a>. Data now runs <strong>1 Jan 1989 to 31 Jul 2026</strong>. The “warmest summer” framing is context from the Met Office (link above); this analysis works only with reservoir <em>levels</em>, not temperature. Notebook: <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-reservoir-levels/reservoir_analysis_2026-07.ipynb">reservoir_analysis_2026-07.ipynb</a>.</p>
</blockquote>

<h2 id="first-just-add-the-days">First, just add the days</h2>

<p>No cleverness here. I extended the sawtooth by two months and shaded 2026 so the newest readings stand out against four decades.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/daily-2026.html" title="London reservoir levels, 1989 to 2026, with 2026 highlighted" loading="lazy" style="height:470px"></iframe>
  <figcaption>Every daily reading, 1989-2026; the 2026 window is shaded. Drag to zoom into the tail.</figcaption>
</figure>

<p>Zoom into the right-hand edge and the shape is clear: after peaking in spring, both lines fall through June and July. The orange Lower Thames line keeps going, sliding to <strong>72%</strong> by 31 July (the Lower Lee ends at <strong>82%</strong>). That’s a real drop, but a real drop isn’t a finding. The daily view can’t tell me whether 72% is alarming or ordinary for late July. For that I need to compare 2026 against its own history.</p>

<h2 id="how-unusual-is-it-draw-the-normal-band">How unusual is it? Draw the normal band</h2>

<p>The honest test is the same one that flagged 2022 in the first post: plot 2026 on top of the <strong>normal band</strong>. The grey envelope is the full day-by-day min-to-max of every <em>other</em> year (1989-2025); the dashed line is the typical year. If 2026 sits inside the band, it’s within normal. If it rides the edge, it isn’t.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/normal-band-2026.html" title="2026 reservoir levels against the 37-year normal band, both groups" loading="lazy" style="height:640px"></iframe>
  <figcaption>2026 (coloured) against the 1989-2025 daily range (grey) and the typical year (dashed), for each group.</figcaption>
</figure>

<p>Here the two groups part ways, and this is the actual story.</p>

<p>The <strong>Lower Thames</strong> (bottom panel) starts the year mid-band, tracks the typical line into spring, then peels away and rides the <strong>bottom edge of the grey band</strong> all summer. Its 72% on 31 July isn’t just low: it <em>matches the lowest level ever recorded for that date</em> in the 37-year record. For late July, the Thames group has never been emptier. It’s on the floor of its own range.</p>

<p>The <strong>Lower Lee</strong> (top panel) tells a softer version. It fell too, but it’s sitting comfortably <em>inside</em> the band, around the 24th percentile for the date: low-ish, nowhere near a record. Same city, same summer, two different responses. The structural gap between these systems, which the first post flagged, is doing real work here.</p>

<p>So the headline the slogan wanted, “record heat empties the reservoirs,” is half right and half wrong, and the interesting half is the wrong one. One group is at a record low for the date; the other is just a bit dry. Which raises the obvious question: <em>did the Thames get there because this summer drained it unusually fast?</em></p>

<h2 id="was-the-drawdown-actually-the-fastest">Was the drawdown actually the fastest?</h2>

<p>A hot, dry summer should show up as <strong>speed</strong>: how far the reservoirs fall between late spring and mid-summer. So for every year I measured the June-to-July drawdown (average level in the last week of May minus the last week of July) and ranked them.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/summer-drawdown.html" title="Steepest May-to-July reservoir drawdowns by year, Lower Thames" loading="lazy" style="height:470px"></iframe>
  <figcaption>Lower Thames: the size of the late-May-to-late-July fall, by year. 2026 is in red.</figcaption>
</figure>

<p>And here the data complicates the neat narrative. The Lower Thames fell about <strong>17 points</strong> this summer, steep enough to rank <strong>5th-fastest</strong> of 38 summers, but not the record. <strong>2022</strong> fell faster (about 24 points), and so did several other years. The Lower Lee’s 13-point fall ranks 7th; its record belongs to <strong>2018</strong> (about 22 points). A summer the Met Office thinks may be the warmest on record did <strong>not</strong> produce the fastest drawdown on record.</p>

<p>So how is the Thames at a record low for the date if the fall wasn’t record-breaking? Because of where it <em>started</em>. The Lee refilled to a brimming 100% over winter; the Thames only reached about 93%. A slightly lower starting line, plus a steep but not unprecedented fall, was enough to land it on the floor. The record isn’t about how hard this summer pulled. It’s about the summer landing on a spring that never quite topped the Thames up.</p>

<h2 id="what-the-data-says">What the data says</h2>

<p>Reading the data rather than the headline:</p>

<ul>
  <li><strong>The Lower Thames is at a record low <em>for the date</em>:</strong> 72% on 31 July, matching the emptiest late-July in 37 years. That part of the scary story is true.</li>
  <li><strong>But the drawdown wasn’t the fastest.</strong> 2022 and 2018 both drained faster. Record heat did not mean record-speed emptying.</li>
  <li><strong>Starting height mattered as much as the heat.</strong> The Thames began summer a touch lower than the Lee and ended on the floor; the Lee started full and stayed inside its normal band.</li>
  <li><strong>The real test is still ahead.</strong> July isn’t the seasonal trough; that comes in September or October. If the Thames keeps hugging the bottom edge, <em>that</em> autumn minimum is the number to watch, against the deep droughts of 1996-97 the first post dug up.</li>
</ul>

<p>None of this downplays the summer. A record-low-for-the-date reservoir is worth watching. But the outcome I’d have <em>guessed</em> (“hottest summer, fastest drain”) isn’t the one the data supports. The reservoirs are spending a full winter’s savings fast, not running on empty, and the number that will really matter hasn’t happened yet. I’ll come back in October. The <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-reservoir-levels/reservoir_analysis_2026-07.ipynb">notebook is here</a> if you want to watch the floor with me.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="London" /><category term="Open Data" /><category term="Plotly" /><category term="Water" /><summary type="html"><![CDATA[Six weeks after the first reservoir post I added June and July 2026, the start of what the Met Office says may be the UK's warmest summer on record. The Lower Thames is now at its lowest level for the date in 37 years, but the drawdown that got it there wasn't the fastest. A short, data-first follow-up.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/london-reservoirs-hot-summer.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/london-reservoirs-hot-summer.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Who London remembers: 1,037 blue plaques and the very small city inside the city</title><link href="https://koulakhilesh.github.io/writing/london-blue-plaques/" rel="alternate" type="text/html" title="Who London remembers: 1,037 blue plaques and the very small city inside the city" /><published>2026-07-11T11:00:00+01:00</published><updated>2026-08-15T11:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/london-blue-plaques</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/london-blue-plaques/"><![CDATA[<div style="text-align:center;margin:2rem 0 2.4rem;">
<svg width="188" height="188" viewBox="0 0 200 200" role="img" aria-label="A stylised blue plaque reading 1,037 blue plaques, analysed here, 2026">
  <defs>
    <path id="arcTop" d="M 42 100 A 58 58 0 0 1 158 100" fill="none" />
    <path id="arcBot" d="M 46 100 A 54 54 0 0 0 154 100" fill="none" />
  </defs>
  <circle cx="100" cy="100" r="94" fill="#1c5fb0" />
  <circle cx="100" cy="100" r="94" fill="none" stroke="#0e3f80" stroke-width="4" />
  <circle cx="100" cy="100" r="78" fill="none" stroke="#ffffff" stroke-width="2" opacity="0.85" />
  <text fill="#fff" font-family="Georgia, serif" font-size="12.5" letter-spacing="2.5" text-anchor="middle">
    <textPath href="#arcTop" startOffset="50%">ANALYSED HERE</textPath>
  </text>
  <text fill="#fff" font-family="Georgia, serif" font-size="12.5" letter-spacing="3" text-anchor="middle">
    <textPath href="#arcBot" startOffset="50%">2026</textPath>
  </text>
  <text x="100" y="90" fill="#fff" font-family="Georgia, serif" font-size="30" font-weight="700" text-anchor="middle">1,037</text>
  <text x="100" y="113" fill="#fff" font-family="Georgia, serif" font-size="14" letter-spacing="3.5" text-anchor="middle">BLUE</text>
  <text x="100" y="131" fill="#fff" font-family="Georgia, serif" font-size="14" letter-spacing="3.5" text-anchor="middle">PLAQUES</text>
</svg>
</div>

<p>There’s a game I play on walks. Spot a blue plaque, cover the name with my thumb, and guess who lived there. I’m almost always wrong, and that’s the point: a blue plaque is a tiny act of civic memory, a decision that <em>this</em> person, in <em>this</em> building, is worth stopping a stranger in the street for. I’d walked past dozens before it occurred to me to ask the obvious question: who makes that list, and what does the whole list look like at once?</p>

<p>English Heritage runs the scheme, and their website will show you the plaques twelve at a time. I wanted all of them on one screen. So I scraped the lot, all <strong>1,037 plaques</strong>, and went looking for the shape of London’s memory. What I found is that the shape itself is the story.</p>

<h2 id="what-the-data-is">What the data is</h2>

<p>Every plaque has a page: a name, an address, a category, the exact words on the ceramic, and a pin on a map. The listing is served by a quiet little search API, and each plaque’s coordinates are tucked into the map code on its own page. Two passes (one for the list, one to open all 1,037 detail pages) and it folds into a single tidy table.</p>

<blockquote>
  <p><strong>Data &amp; licence.</strong> Plaque data scraped in August 2026 from the <a href="https://www.english-heritage.org.uk/visit/blue-plaques/">English Heritage blue plaques</a> site (listing API + individual plaque pages). The content belongs to English Heritage; this is a personal, non-commercial analysis of publicly visible information, with attribution. The full scraper and analysis notebook live in my <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-blue-plaques/blue_plaques_analysis.ipynb">CodePlayground repo</a>.</p>
</blockquote>

<p>Before trusting any of it, I checked how complete each field was. Names, boroughs, coordinates, inscriptions and categories all came back at basically 100%. Birth and death years covered 96% of entries: 991 of the plaques commemorate a specific <em>person</em> (the rest mark buildings and events). One thing I <em>couldn’t</em> get: the date each plaque went up. It isn’t published anywhere on the site, which quietly kills the question I most wanted to ask: how long after death does recognition arrive? Some questions the data just won’t answer, and it’s more honest to say so than to fudge it.</p>

<h2 id="first-put-them-all-on-the-map">First, put them all on the map</h2>

<p>No aggregation, no cleverness. Just drop all 1,036 plaques-with-coordinates onto London and colour them by category.</p>

<figure class="chart">
  <iframe src="/assets/plaques/plaque-map.html" title="Every London blue plaque, mapped and coloured by category" loading="lazy" style="height:640px"></iframe>
  <figcaption>All 1,036 plaques with coordinates. Hover any dot for the name, borough and category; drag to pan, scroll to zoom.</figcaption>
</figure>

<p>You don’t need statistics to see it. There’s a dense, glowing core in the centre and west, and then the rest of London, enormous, populous, historic London, thins out to scattered dots. My first thought was that I’d made a mistake. I hadn’t. London’s official memory really is packed into a small footprint. <em>(Just how tightly packed, statistically, is the subject of <a href="#a-companion-piece">the sequel to this post</a>, where the same points become a geometry problem.)</em></p>

<h2 id="three-boroughs-hold-two-thirds-of-it">Three boroughs hold two-thirds of it</h2>

<p>I counted, and the concentration is even starker than the map suggests.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value">69%</div>
    <div class="stat-label">in just 3 boroughs</div>
  </div>
  <div class="stat">
    <div class="stat-value">31<span class="stat-unit">/33</span></div>
    <div class="stat-label">boroughs with any plaque</div>
  </div>
  <div class="stat">
    <div class="stat-value">10</div>
    <div class="stat-label">on Cheyne Walk alone</div>
  </div>
</div>

<p><strong>Just three boroughs, Westminster, Kensington &amp; Chelsea, and Camden, hold 69% of every blue plaque in London.</strong> Westminster alone has 334, nearly a third of the whole scheme. Meanwhile Barking &amp; Dagenham, Sutton and the City of London manage one apiece.</p>

<figure class="chart">
  <iframe src="/assets/plaques/by-borough.html" title="Blue plaques by London borough" loading="lazy" style="height:520px"></iframe>
  <figcaption>Plaques per borough (top 15). The top three dwarf everything else.</figcaption>
</figure>

<p>Zoom in past the borough line and the clustering gets almost comically specific. The single most-plaqued street in London is <strong>Cheyne Walk</strong> in Chelsea, with ten; the densest postcode district is <strong>NW3, Hampstead</strong>, with sixty-nine.</p>

<figure class="chart">
  <iframe src="/assets/plaques/top-streets.html" title="London&#39;s most-plaqued streets" loading="lazy" style="height:520px"></iframe>
  <figcaption>The most-plaqued streets. Bedford Square, Gower Street and Queen Anne's Gate follow Chelsea's riverside.</figcaption>
</figure>

<p>And each borough remembers a <em>different kind</em> of person. Westminster’s plaques lean towards <strong>politicians</strong>; Kensington &amp; Chelsea and Camden lean towards <strong>writers and artists</strong>. Westminster remembers power; Chelsea remembers art. That’s not a coincidence: it’s the first hint of the feedback loop I’ll come back to.</p>

<p>Some of this is honest history: the West End and Bloomsbury <em>were</em> where the writers, scientists and statesmen of a certain era clustered, near the salons and institutions and each other. But some of it is self-fulfilling. Plaques mark the grand, surviving townhouses of central London, and grand surviving townhouses are exactly where well-documented, well-connected, plaque-worthy lives happened. The map isn’t just showing where remarkable people lived. It’s showing where the <em>kind</em> of person the scheme was built to remember lived.</p>

<h2 id="who-gets-remembered">Who gets remembered</h2>

<p>Sort the plaques by what the person is remembered <em>for</em>, and London reveals itself as, above all, a city of writers.</p>

<figure class="chart">
  <iframe src="/assets/plaques/by-category.html" title="Blue plaques by primary category" loading="lazy" style="height:520px"></iframe>
  <figcaption>Primary category per plaque (top 15). Literature leads, ahead of politics, fine arts and music.</figcaption>
</figure>

<p>Literature comes first, then politics and administration, then the fine arts and music. It’s a portrait of what a culture chooses to enshrine: the people who left <em>documents</em> (books, laws, paintings, scores), the kind of legacy that keeps a name legible a century later.</p>

<h2 id="the-gap-isnt-flat">The gap isn’t flat</h2>

<p>Here’s the number the scheme itself is quietly self-conscious about. There’s no gender field in the data, so I inferred it from first names (with an offline name-to-gender library) and honorifics like <em>Sir</em> and <em>Dame</em>. It’s imperfect: 122 names it couldn’t resolve, and it will misfire on some, so I only report the share among names it <em>could</em> resolve. Even so, the result is stark: <strong>fewer than one in five commemorated people are women.</strong></p>

<p>But the single figure hides the interesting part. Split the women’s share by <em>field</em> and it swings wildly, from near-parity in one category to an absolute, unbroken zero in others.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value stat-value--alt">17.5%</div>
    <div class="stat-label">women overall</div>
  </div>
  <div class="stat">
    <div class="stat-value stat-value--alt">0%</div>
    <div class="stat-label">women in engineering, industry &amp; invention</div>
  </div>
  <div class="stat">
    <div class="stat-value">174<span class="stat-unit">:25</span></div>
    <div class="stat-label">Sirs to Dames</div>
  </div>
</div>

<figure class="chart">
  <iframe src="/assets/plaques/female-by-field.html" title="Share of women commemorated, by field" loading="lazy" style="height:540px"></iframe>
  <figcaption>Female share by category (fields with at least 15 plaques). The dashed line is the 17.5% overall average.</figcaption>
</figure>

<p>Women appear most in <strong>philanthropy and reform</strong> (close to half) and on the <strong>stage</strong>: theatre, film, dance. They all but vanish from <strong>politics</strong> (3%) and hit a flat <strong>zero</strong> in engineering, industry and invention. The pattern isn’t really about the plaques; it’s about which doors were open to women in the first place, preserved in ceramic. The honorifics say the same thing more bluntly: 174 knighted <em>Sirs</em> to 25 <em>Dames</em>.</p>

<p>There is one hopeful thread. Stack the commemorated people by their birth decade and the pink band, however thin, grows as you move toward the present.</p>

<figure class="chart">
  <iframe src="/assets/plaques/gender-by-decade.html" title="Commemorated people by birth decade and inferred gender" loading="lazy" style="height:460px"></iframe>
  <figcaption>People commemorated, stacked by birth decade and inferred gender. Watch the recent decades.</figcaption>
</figure>

<p>Among people born in the mid-1700s the female share was under 5%; among those born around 1900 it’s up past 30%. The scheme is trying to correct, and English Heritage has said as much publicly. You can watch that intention arrive, one decade at a time, but you’re watching it climb out of a very deep hole.</p>

<h2 id="what-the-walls-actually-say">What the walls actually say</h2>

<p>The inscriptions are their own small corpus. Throw all 1,037 into a pile and count the words, and the most common ones read like a poem about the dataset in miniature: <em>poet, writer, painter, novelist, artist</em>, and, sitting right among them, <em>Sir</em>.</p>

<figure class="chart">
  <iframe src="/assets/plaques/inscription-words.html" title="Most common words in blue plaque inscriptions" loading="lazy" style="height:560px"></iframe>
  <figcaption>The most common words across all 1,037 inscriptions, filler words removed.</figcaption>
</figure>

<p>There’s also a quiet grammar to how each plaque relates its person to its building. Nearly every one states a verb, and one verb runs away with it.</p>

<figure class="chart">
  <iframe src="/assets/plaques/inscription-verbs.html" title="What the plaques say the person did at the address" loading="lazy" style="height:430px"></iframe>
  <figcaption>Relationship to the address, counted across all inscriptions. "Lived here" appears on 725 of them.</figcaption>
</figure>

<p><strong>725 plaques, seven in ten, simply say “lived here.”</strong> Not born, not died, not worked. Lived. The blue plaque is a marker of <em>domesticity made historic</em>: this ordinary front door held an extraordinary ordinary life. And almost all of them say it on the same object: 88% of the plaques are the familiar ceramic roundel, with a stubborn handful in bronze, stone or slate.</p>

<h2 id="how-long-did-they-live">How long did they live?</h2>

<p>A small, humane detour. For the 991 people, birth and death years give a distribution of lifespans.</p>

<figure class="chart">
  <iframe src="/assets/plaques/lifespans.html" title="Age at death of the commemorated" loading="lazy" style="height:430px"></iframe>
  <figcaption>Age at death, filtered to plausible values. The median is 72.</figcaption>
</figure>

<p>The median is <strong>72</strong>, with a long tail into the nineties: Sir Robert Mayer made it to 106. Which makes a certain grim sense: the surest route onto a wall of long-term public memory is to do enough, for long enough, to be remembered. But the left tail is where the heartbreak lives: a cluster who died in their twenties and got a plaque anyway: John Keats at 26, the sculptor Gaudier-Brzeska at 24, the SOE agent Violette Szabo at 24. Fame usually rewards patience. Just occasionally it rewards a comet.</p>

<h2 id="memory-isnt-merit">Memory isn’t merit</h2>

<p>Put the threads together and they all point the same way. The geography (a tiny central core), the gender (fewer than one in five, and zero in whole professions), the categories (writers and statesmen), the honorifics (Sirs outnumbering Dames seven to one), these aren’t four separate findings. They’re four views of a single filter.</p>

<p>Blue plaques don’t really mark where <em>remarkable</em> people lived. They mark where the <strong>documented, connected, establishment</strong> class lived, and that class was central, male, literary and titled. The filter is self-reinforcing: grand surviving houses in Zone 1 are both where those lives happened and where a plaque can be hung today, so the memory keeps pooling in the same square mile.</p>

<p>None of this is a knock on the plaques. I still play the guessing game, and I still lose. But a wall of memory is also a mirror, and it’s worth occasionally asking who’s reflected in it and who isn’t: the outer boroughs, the women who never got the door opened, the engineers who built the city and got nothing.</p>

<h3 id="a-companion-piece">A companion piece</h3>

<p>That’s the <em>human</em> story. There’s also a purely <em>mathematical</em> one hiding in the same 1,036 dots: how clustered London’s memory is (provably), what territory each plaque owns, and the shortest possible walking tour of the whole city. I pulled that apart in a second post: <strong><a href="/writing/the-geometry-of-londons-blue-plaques/">The geometry of memory →</a></strong>. If you’d rather poke at the data yourself, the <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-blue-plaques/blue_plaques_analysis.ipynb">scraper and notebook</a> are one afternoon of pandas and Plotly. Go find the plaque nobody put up.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="London" /><category term="History" /><category term="Plotly" /><category term="Web Scraping" /><summary type="html"><![CDATA[I scraped every English Heritage blue plaque in London to ask who gets remembered and where. The answer is a surprisingly tiny, surprisingly male, surprisingly literary corner of the map, and the shape of it says something uncomfortable.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/london-blue-plaques.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/london-blue-plaques.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The geometry of London’s blue plaques: Voronoi, clustering, and a travelling-salesman tour</title><link href="https://koulakhilesh.github.io/writing/the-geometry-of-londons-blue-plaques/" rel="alternate" type="text/html" title="The geometry of London’s blue plaques: Voronoi, clustering, and a travelling-salesman tour" /><published>2026-07-11T10:00:00+01:00</published><updated>2026-08-15T10:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/the-geometry-of-londons-blue-plaques</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/the-geometry-of-londons-blue-plaques/"><![CDATA[<div style="text-align:center;margin:2rem 0 2.4rem;">
<svg width="188" height="188" viewBox="0 0 200 200" role="img" aria-label="A stylised blue plaque reading 1,036 points, the geometry of memory, 2026">
  <defs>
    <path id="arcTop2" d="M 36 100 A 64 64 0 0 1 164 100" fill="none" />
    <path id="arcBot2" d="M 46 100 A 54 54 0 0 0 154 100" fill="none" />
  </defs>
  <circle cx="100" cy="100" r="94" fill="#1c5fb0" />
  <circle cx="100" cy="100" r="94" fill="none" stroke="#0e3f80" stroke-width="4" />
  <circle cx="100" cy="100" r="78" fill="none" stroke="#ffffff" stroke-width="2" opacity="0.85" />
  <text fill="#fff" font-family="Georgia, serif" font-size="10" letter-spacing="0.8" text-anchor="middle">
    <textPath href="#arcTop2" startOffset="50%">THE GEOMETRY OF MEMORY</textPath>
  </text>
  <text fill="#fff" font-family="Georgia, serif" font-size="12.5" letter-spacing="3" text-anchor="middle">
    <textPath href="#arcBot2" startOffset="50%">2026</textPath>
  </text>
  <text x="100" y="90" fill="#fff" font-family="Georgia, serif" font-size="30" font-weight="700" text-anchor="middle">1,036</text>
  <text x="100" y="113" fill="#fff" font-family="Georgia, serif" font-size="14" letter-spacing="3.5" text-anchor="middle">POINTS</text>
  <text x="100" y="131" fill="#fff" font-family="Georgia, serif" font-size="14" letter-spacing="3.5" text-anchor="middle">ON A MAP</text>
</svg>
</div>

<p>In <a href="/writing/london-blue-plaques/">the first half of this</a> I scraped every English Heritage blue plaque in London and asked <em>who</em> gets remembered. The headline was that London’s memory is astonishingly concentrated: you could see it just by dropping the dots on a map.</p>

<p>But “you can see it” is not the same as “it’s true.” This post is the data scientist’s follow-up: take the same 1,036 geolocated plaques and treat them as a <strong>spatial point pattern</strong> and an <strong>optimisation problem</strong>. Can I <em>prove</em> the clustering? What territory does each plaque own? Where are the hotspots, found by an algorithm, not my eyeballs? And, for fun, what’s the shortest walking tour of the whole city?</p>

<blockquote>
  <p><strong>Method note.</strong> The 1,036 plaque locations were scraped in August 2026 from the <a href="https://www.english-heritage.org.uk/visit/blue-plaques/">English Heritage blue plaques</a> site, as a personal, non-commercial analysis with attribution (the full licence note is in <a href="/writing/london-blue-plaques/">part one</a>). Everything below works in the <strong>British National Grid (EPSG:27700)</strong>, so distances and areas are in real metres, not degrees. Code lives in the <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-blue-plaques/blue_plaques_geospatial.ipynb">geospatial notebook</a>.</p>
</blockquote>

<h2 id="is-it-clustered-prove-it">Is it clustered? Prove it.</h2>

<p>Before any picture, a statistic. The eye is easily fooled, so the honest first move is a test. The <strong>Clark-Evans index</strong> compares the average distance from each plaque to its <em>nearest</em> neighbour against what you’d expect if the same number of plaques were scattered at random across the same area. Below 1 means clustered; above 1 means spread out; a big <em>z</em>-score means it’s not luck.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value">0.54</div>
    <div class="stat-label">Clark-Evans R (1 = random)</div>
  </div>
  <div class="stat">
    <div class="stat-value">-28</div>
    <div class="stat-label">z-score (|z|&gt;2 is significant)</div>
  </div>
  <div class="stat">
    <div class="stat-value">100<span class="stat-unit"> m</span></div>
    <div class="stat-label">median nearest neighbour</div>
  </div>
</div>

<p>The index comes out at <strong>R = 0.54</strong>, with a <em>z</em>-score of about <strong>-28</strong>. In plain terms: plaques sit roughly <em>half</em> as far from their nearest neighbour as random scattering would predict, and the chance of that happening by luck is essentially nil. The median plaque has another plaque just <strong>100 metres</strong> away. Here’s the distribution behind the number:</p>

<figure class="chart">
  <iframe src="/assets/plaques/nn-distances.html" title="Distribution of nearest-neighbour distances between plaques" loading="lazy" style="height:400px"></iframe>
  <figcaption>Distance from each plaque to its nearest neighbour. The spike near zero is the signature of clustering.</figcaption>
</figure>

<h2 id="your-nearest-plaque-a-voronoi-map">Your nearest plaque: a Voronoi map</h2>

<p>Statistics proven, now the pretty part. A <strong>Voronoi diagram</strong> carves the plane into one cell per plaque, where every point in a cell is closer to <em>that</em> plaque than to any other, the plaque’s “catchment area.” Colour each cell by its size and the density story becomes visceral.</p>

<figure class="chart">
  <iframe src="/assets/plaques/voronoi.html" title="Voronoi diagram of London&#39;s blue plaques, coloured by cell area" loading="lazy" style="height:640px"></iframe>
  <figcaption>Each cell is the territory of one plaque; darker means smaller. The centre is a mosaic of tiny tiles; the edges are single plaques owning whole boroughs.</figcaption>
</figure>

<p>In the West End the cells are so small they blur into a mosaic, a plaque every hundred metres, each owning a scrap of pavement. Out toward Bromley or Croydon a lone plaque can own kilometres in every direction. The <em>area</em> of your nearest-plaque territory is, in effect, an inverse density map, and it screams the same thing the dots did, now with an area attached to it.</p>

<p><em>Try it: <a href="/lab/#plaque-voronoi">break the clustering yourself in the Lab’s Voronoi playground →</a></em></p>

<h2 id="the-heat-of-memory">The heat of memory</h2>

<p>The same information, smoothed into a continuous surface: a kernel-style density heatmap. No borough lines, no cells, just where memory glows hottest.</p>

<figure class="chart">
  <iframe src="/assets/plaques/density.html" title="Density heatmap of London&#39;s blue plaques" loading="lazy" class="chart--dark" style="height:620px"></iframe>
  <figcaption>A density surface over the plaques. One bright ridge runs from Bloomsbury through Mayfair to Chelsea.</figcaption>
</figure>

<h2 id="hotspots-found-by-algorithm">Hotspots, found by algorithm</h2>

<p>I don’t want to <em>decide</em> where the clusters are: that’s cheating. So I handed the job to <strong>DBSCAN</strong>, a density-based clustering algorithm: any group of at least six plaques all within 350 metres of one another becomes a cluster; everything else is “scattered.” Told only the coordinates, it recovers the hotspots on its own.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value">14</div>
    <div class="stat-label">hotspots discovered</div>
  </div>
  <div class="stat">
    <div class="stat-value">287</div>
    <div class="stat-label">plaques in the biggest (Westminster)</div>
  </div>
  <div class="stat">
    <div class="stat-value">232</div>
    <div class="stat-label">in the second (Kensington)</div>
  </div>
</div>

<figure class="chart">
  <iframe src="/assets/plaques/hotspots.html" title="DBSCAN clustering of blue plaques into hotspots" loading="lazy" style="height:640px"></iframe>
  <figcaption>Blue dots belong to a hotspot; grey dots are scattered. DBSCAN found 14 clusters knowing only the coordinates.</figcaption>
</figure>

<p>It lands exactly where you’d expect: one giant blob over Westminster, another over Kensington &amp; Chelsea, satellites in Bloomsbury and Hampstead, but the point is that it <em>found</em> them. Half of all London’s plaques fall into just the two largest clusters.</p>

<h2 id="the-memory-network">The memory network</h2>

<p>A different lens from graph theory: what’s the shortest possible set of links that connects <em>every</em> plaque into one network, with no loops? That’s a <strong>minimum spanning tree</strong>, and drawn on the map it looks like the nervous system of London’s memory.</p>

<figure class="chart">
  <iframe src="/assets/plaques/memory-network.html" title="Minimum spanning tree connecting all blue plaques" loading="lazy" style="height:640px"></iframe>
  <figcaption>The minimum spanning tree over all 1,036 plaques: 371 km of shortest-possible links.</figcaption>
</figure>

<p>The whole tree is <strong>371 km</strong> long, but look at how the “wire” is spent. It’s dense and short in the centre, where neighbours are 100 m apart, and it throws long lonely spans out to the isolated plaques on the fringe.</p>

<h2 id="the-grand-tour-a-travelling-salesman-in-london">The Grand Tour: a travelling salesman in London</h2>

<p>Finally, the classic. Pick one representative plaque per borough (the one nearest each borough’s centroid) and ask: what’s the shortest loop that visits all of them and returns home? That’s the <strong>Travelling Salesman Problem</strong>, and I solved it from scratch rather than calling a solver, because the <em>method</em> is half the fun.</p>

<p>The recipe is two moves. First, a <strong>nearest-neighbour</strong> heuristic: start somewhere, always walk to the closest unvisited borough. It’s greedy and it leaves ugly crossings. Then <strong>2-opt</strong>: repeatedly find two edges that cross, snip them, and reconnect the other way, which always shortens the tour, until no swap helps.</p>

<div class="stats">
  <div class="stat">
    <div class="stat-value stat-value--alt">219<span class="stat-unit"> km</span></div>
    <div class="stat-label">greedy nearest-neighbour tour</div>
  </div>
  <div class="stat">
    <div class="stat-value">180<span class="stat-unit"> km</span></div>
    <div class="stat-label">after 2-opt</div>
  </div>
  <div class="stat">
    <div class="stat-value">18%</div>
    <div class="stat-label">shorter, for a few lines of code</div>
  </div>
</div>

<figure class="chart">
  <iframe src="/assets/plaques/grand-tour.html" title="Optimised travelling-salesman tour through all 31 boroughs" loading="lazy" style="height:640px"></iframe>
  <figcaption>The optimised loop through one plaque per borough: 180 km, no crossings. Hover a node for its borough.</figcaption>
</figure>

<p>Nearest-neighbour alone gives a <strong>219 km</strong> tour with tell-tale crossings. A pass of 2-opt untangles it down to <strong>180 km</strong>, an 18% saving for a couple of dozen lines of Python, and the visual proof is that the crossings are gone. That is the whole point of local search: a dumb first guess, plus a simple “is this knot removable?” rule, gets you most of the way to optimal.</p>

<h2 id="what-the-geometry-taught-me">What the geometry taught me</h2>

<ul>
  <li>The clustering isn’t a trick of the eye: <strong>Clark-Evans R = 0.54, z of about -28</strong> makes it statistically undeniable.</li>
  <li><strong>Voronoi territories</strong> turn density into area: tiny tiles downtown, whole boroughs at the edge.</li>
  <li><strong>DBSCAN</strong> rediscovers the hotspots from coordinates alone, a nice reminder that the structure is <em>in the data</em>, not in my assumptions.</li>
  <li>A hand-rolled <strong>nearest-neighbour + 2-opt</strong> TSP shaves ~18% off the naive route, which is the whole value proposition of local-search optimisation in miniature.</li>
</ul>

<p>If the <a href="/writing/london-blue-plaques/">first post</a> was about <em>who</em> London remembers, this one was about the <em>shape</em> of that memory, and the shape, it turns out, is provably, beautifully lopsided. The <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-blue-plaques/blue_plaques_geospatial.ipynb">notebook is here</a> if you want to run the tour yourself.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="Geospatial" /><category term="Optimization" /><category term="Plotly" /><category term="London" /><summary type="html"><![CDATA[A companion piece that treats 1,036 blue plaques as a spatial point pattern and an optimisation problem: proving the clustering with statistics, carving London into nearest-plaque territories, and hand-rolling a travelling-salesman tour of every borough.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/the-geometry-of-londons-blue-plaques.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/the-geometry-of-londons-blue-plaques.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">London’s reservoirs have a heartbeat (and the drought you remember isn’t the one that ran them driest)</title><link href="https://koulakhilesh.github.io/writing/london-reservoirs/" rel="alternate" type="text/html" title="London’s reservoirs have a heartbeat (and the drought you remember isn’t the one that ran them driest)" /><published>2026-07-04T00:00:00+01:00</published><updated>2026-08-15T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/london-reservoirs</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/london-reservoirs/"><![CDATA[<p>I went looking for the summer of 2022, a summer I never saw.</p>

<p>I landed in London that November, a few months too late for it. You don’t have to have been here to have heard about it, though. The grass had gone the colour of straw, the hosepipe bans had come and gone, and “drought” was a headline word all summer. It was one of the first pieces of local weather-lore I picked up. So when I found a dataset with <strong>daily</strong> reservoir levels for London going back to 1989, the first thing I did was hunt for 2022, the drought everyone had told me about. I expected a cliff. The worst dip in the whole record.</p>

<p>It wasn’t. And that little wrongness, a memory I’d inherited rather than lived, is what turned a quick afternoon poke into this post.</p>

<h2 id="what-the-data-is">What the data is</h2>

<p>The <a href="https://data.london.gov.uk/dataset/london-reservoir-levels-24ry5">London Datastore</a> publishes the level of two big reservoir groups that supply the city:</p>

<ul>
  <li>the <strong>Lower Lee</strong> group, and</li>
  <li>the <strong>Lower Thames</strong> group.</li>
</ul>

<p>One row per day, one number per group, all the way from <strong>1 January 1989</strong> to <strong>31 May 2026</strong>. Each number is just <em>percent of capacity</em>: how full the reservoirs were that day. That’s the entire dataset. No fancy features, no engineered columns. Just ~13,665 days of “how much water is in the tank.”</p>

<p>The part that delighted me, before I plotted anything, was how <em>complete</em> it is. 37 years, one reading every single day, and when I checked for gaps in the calendar there were <strong>none</strong>. 16 missing values in total, out of about 27,000. Public datasets tend to be messier than that. This one is a quiet feat of record-keeping.</p>

<blockquote>
  <p><strong>Data &amp; licence.</strong> London reservoir levels via the <a href="https://data.london.gov.uk/dataset/london-reservoir-levels-24ry5">London Datastore</a>, sourced from the Environment Agency’s <a href="https://www.gov.uk/government/collections/water-situation-reports-for-england">water situation reports for England</a>. Contains public sector information licensed under the <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/2/">Open Government Licence v2.0</a>. The full analysis notebook lives in my <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-reservoir-levels/reservoir_analysis.ipynb">CodePlayground repo</a>.</p>
</blockquote>

<h2 id="where-these-two-groups-sit">Where these two groups sit</h2>

<p>The dataset never says <em>where</em> these reservoirs are; “Lower Lee” and “Lower Thames” are just two column names. They’re real bodies of water, sitting in two very different corners of London. The <strong>Lower Lee</strong> group is the Lee Valley chain up in the north-east, above Tottenham and Walthamstow. The <strong>Lower Thames</strong> group is the cluster of big reservoirs out west, towards Staines and Heathrow. Two systems, one city, fed by separate rivers.</p>

<p>To draw them, I pulled the reservoir outlines straight from OpenStreetMap and coloured each polygon by its group. There are no coordinates in the levels CSV, so this is a separate open-data layer stitched on top.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/reservoir-map.html" title="Map of the Lower Lee and Lower Thames reservoir groups" loading="lazy" style="height:520px"></iframe>
  <figcaption>The two groups sit in opposite corners of London: the Lower Lee (green) in the north-east, the Lower Thames (orange) out to the south-west. Hover any reservoir for its name.</figcaption>
</figure>

<blockquote>
  <p><strong>Map data.</strong> Reservoir outlines and basemap © <a href="https://www.openstreetmap.org/copyright">OpenStreetMap</a> contributors, licensed under the Open Database Licence (ODbL). Fetched via the Overpass API in the analysis notebook.</p>
</blockquote>

<h2 id="first-just-plot-everything">First, just plot everything</h2>

<p>Before slicing or summarising anything, it pays to be a little dumb and just plot <em>every single day</em>. No aggregation, no smoothing. Just drop all 13,665 points on a chart and see what the shape tells you.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/daily-levels.html" title="London reservoir levels, daily, 1989 to 2026" loading="lazy" style="height:470px"></iframe>
  <figcaption>Every daily reading, 1989-2026. Drag to zoom into a year, double-click to reset.</figcaption>
</figure>

<p>There it is: a <strong>sawtooth</strong>. Up through the winter, down through the summer, over and over for close to four decades. Your eye catches the deep dips: the mid-1990s, a rough patch around 2005-06, and yes, 2022. Those are the droughts. But hold that thought about 2022, because the daily view is too noisy to rank them. For that we need to fold time.</p>

<h2 id="the-heartbeat-what-an-average-year-looks-like">The heartbeat: what an average year looks like</h2>

<p>If every year follows near enough the same fill-and-drain rhythm, then I should be able to squash all 37 years onto a single 12-month clock and see the underlying pulse. Statisticians call this a <em>climatology</em>. I like to think of it as taking the reservoir’s resting heart rate.</p>

<p>So I averaged every January together, every February together, and so on. The shaded bands show the full range each month has ever spanned; the solid lines are the averages:</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/average-year.html" title="The average London water year, by month" loading="lazy" style="height:460px"></iframe>
  <figcaption>All 37 years folded onto one 12-month clock. Hover any month to compare the two groups.</figcaption>
</figure>

<p>The Lower Thames peaks in <strong>late winter</strong> (think February, when it’s been raining for months and nobody’s watering a garden) and bottoms out around <strong>September to October</strong>, right before the autumn rains kick in and start the refill. The Lower Lee does the same dance but sits a few points lower the whole way round. That gap between the two groups isn’t random jitter; it’s there in almost every year. Two reservoir systems, same city, different personalities.</p>

<h2 id="which-years-ran-driest">Which years ran driest?</h2>

<p>Here’s where I went back for 2022. The fairest single number for “how bad was that year” is the <strong>annual minimum</strong>: the lowest the reservoirs fell at the worst moment of that year. Short bars are the scary years.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/annual-minimum.html" title="Lowest reservoir level reached each year" loading="lazy" style="height:460px"></iframe>
  <figcaption>The lowest point each year. Hover for exact values; 2026 is a partial year (to May).</figcaption>
</figure>

<p>This is the bit that made me sit up. I’d arrived to the <em>story</em> of 2022 rather than the summer itself, so I had no yardstick of my own to check it against. The numbers were my only witness. For the <strong>Lower Thames</strong>, the driest years on record are:</p>

<table>
  <thead>
    <tr>
      <th>Year</th>
      <th style="text-align: right">Lowest level</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>1996</strong></td>
      <td style="text-align: right">42%</td>
    </tr>
    <tr>
      <td><strong>1997</strong></td>
      <td style="text-align: right">44%</td>
    </tr>
    <tr>
      <td>2003</td>
      <td style="text-align: right">46%</td>
    </tr>
    <tr>
      <td>1990</td>
      <td style="text-align: right">48%</td>
    </tr>
    <tr>
      <td>2022</td>
      <td style="text-align: right">50%</td>
    </tr>
  </tbody>
</table>

<p>2022, the drought I’d been assured was the worst, comes in <strong>fifth</strong>. The brutal squeeze was <strong>1996-97</strong>, a slow two-year drawdown that pulled the Thames group down to 42%. For the Lower Lee, the deepest year was <strong>1991</strong> (48%), with 1992 close behind. The droughts that ran London’s reservoirs driest happened in the <em>early-to-mid nineties</em>, and most of us have forgotten them.</p>

<p>Memory is a headline. The data is a diary. When they disagree, I know which one I trust.</p>

<h2 id="putting-2022-under-the-microscope">Putting 2022 under the microscope</h2>

<p>None of this means 2022 wasn’t a real drought. It was. “Real drought” and “worst on record” are two different claims, though, and the honest way to tell them apart is to draw the year against its own history. Here’s 2022 for the Lower Thames, plotted on top of the <strong>normal band</strong>: the grey envelope is the full min-to-max range every <em>other</em> year has ever traced, day by day, and the dashed line is the typical year.</p>

<figure class="chart">
  <iframe src="/assets/reservoirs/drought-2022-band.html" title="2022 reservoir levels versus the 37-year normal band" loading="lazy" style="height:470px"></iframe>
  <figcaption>2022 (orange) against the 1989-2026 daily range (grey band) and the typical year (dashed).</figcaption>
</figure>

<p>You can watch the story unfold across the year. 2022 starts out ordinary, then peels away from the typical line through that hot, dry summer and rides along the <strong>bottom edge</strong> of the band by autumn, low for the time of year, hugging the record floor without breaking it. That’s what a bad year looks like when it isn’t the record. The normal band is, to me, the single most honest chart in the whole notebook, because it answers the one question that matters: <em>how unusual was this?</em></p>

<h2 id="what-i-took-away">What I took away</h2>

<ul>
  <li><strong>Water has a heartbeat.</strong> Fill in winter, draw down to an autumn low, then do it again the next year. Once you see the rhythm, it shows up everywhere.</li>
  <li><strong>The Lower Thames runs fuller than the Lower Lee.</strong> Same pattern, year after year. It’s baked into how the two systems work.</li>
  <li><strong>The drought you remember isn’t the worst one.</strong> 2022 made the news. 1996-97 set the record. The loud year sticks; the deep years fade.</li>
  <li><strong>Complete data is a gift.</strong> The cleaning step was dull, so every surprise waited in the findings instead.</li>
</ul>

<p>If you want to poke at it yourself (different reservoir, different year, your own definition of “drought”), the whole thing is a single <a href="https://github.com/koulakhilesh/CodePlayground/blob/main/london-reservoir-levels/reservoir_analysis.ipynb">Jupyter notebook</a> with pandas and Plotly. Grab the CSV from the London Datastore, drop it in, and go looking for your own wrong memory. That’s the fun part.</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="Data Science" /><category term="London" /><category term="Open Data" /><category term="Plotly" /><category term="Water" /><summary type="html"><![CDATA[37 years of daily reservoir readings, one chart at a time, and a small surprise hiding in the numbers everyone thinks they remember.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://koulakhilesh.github.io/assets/social/london-reservoirs.png" /><media:content medium="image" url="https://koulakhilesh.github.io/assets/social/london-reservoirs.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Understanding the Battery Module in OpenEnergy</title><link href="https://koulakhilesh.github.io/writing/understanding-the-battery/" rel="alternate" type="text/html" title="Understanding the Battery Module in OpenEnergy" /><published>2024-06-05T00:00:00+01:00</published><updated>2024-06-05T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/understanding-the-battery</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/understanding-the-battery/"><![CDATA[<p>Hello everyone! Today, we’re going to dive into the <code class="language-plaintext highlighter-rouge">battery.py</code> module of our OpenEnergy project. This module is the heart of our energy storage system simulation, and it’s where all the magic happens. You can find the complete code in our <a href="https://github.com/koulakhilesh/OpenEnergy/">GitHub repository</a>. This module is a great example of how to model a battery’s behavior in Python.</p>

<h2 id="overview">Overview</h2>

<p>The <code class="language-plaintext highlighter-rouge">battery.py</code> module contains several classes that model different aspects of a battery’s behavior:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">BasicSOHCalculator</code>: Calculates the battery’s State of Health (SOH) based on the energy cycled and depth of discharge (DOD).</li>
  <li><code class="language-plaintext highlighter-rouge">TemperatureEfficiencyAdjuster</code>: Adjusts the battery’s charging and discharging efficiencies based on temperature.</li>
  <li><code class="language-plaintext highlighter-rouge">Battery</code>: Represents a battery with various properties and methods for charging and discharging.</li>
</ul>

<p>Let’s dive into each of these classes and understand how they work.</p>

<h2 id="basicsohcalculator">BasicSOHCalculator</h2>

<p>This class calculates the battery’s State of Health (SOH) based on the energy cycled and depth of discharge (DOD). The <code class="language-plaintext highlighter-rouge">calculate_soh</code> method takes in the initial SOH, the amount of energy cycled, and the DOD, and calculates the new SOH based on a degradation rate.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">calculate_soh</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">soh</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span> <span class="n">energy_cycled_mwh</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span> <span class="n">dod</span><span class="p">:</span> <span class="nb">float</span><span class="p">):</span>
    <span class="s">"""
    Calculates the State of Health (SOH) based on the given parameters.

    Args:
        soh (float): The initial State of Health.
        energy_cycled_mwh (float): The amount of energy cycled in megawatt-hours.
        dod (float): The depth of discharge as a fraction (0 to 1).

    Returns:
        float: The updated State of Health (SOH) after degradation.

    """</span>
    <span class="n">base_degradation</span> <span class="o">=</span> <span class="mf">0.000005</span>
    <span class="n">dod_factor</span> <span class="o">=</span> <span class="mi">2</span> <span class="k">if</span> <span class="n">dod</span> <span class="o">&gt;</span> <span class="mf">0.5</span> <span class="k">else</span> <span class="mi">1</span>
    <span class="n">degradation_rate</span> <span class="o">=</span> <span class="n">base_degradation</span> <span class="o">*</span> <span class="n">energy_cycled_mwh</span> <span class="o">*</span> <span class="n">dod_factor</span>
    <span class="k">return</span> <span class="n">soh</span> <span class="o">*</span> <span class="p">(</span><span class="mi">1</span> <span class="o">-</span> <span class="n">degradation_rate</span><span class="p">)</span>
</code></pre></div></div>

<p>In the accompanying graph, observe the progression of State of Charge (SOC) and State of Health (SOH) values over time, illustrating the dynamic behavior of the battery system.
<img src="https://raw.githubusercontent.com/koulakhilesh/OpenEnergy/master/images/notebook/assets/battery_simulation.png" alt="SOC and SOH" /></p>

<h2 id="temperatureefficiencyadjuster">TemperatureEfficiencyAdjuster</h2>

<p>This class adjusts the battery’s charging and discharging efficiencies based on the temperature. The <code class="language-plaintext highlighter-rouge">adjust_efficiency</code> method takes in the current temperature and the current efficiencies, and adjusts them based on a simple rule: for every degree Celsius the temperature is away from 25°C, the efficiencies decrease by 1%.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">adjust_efficiency</span><span class="p">(</span>
    <span class="bp">self</span><span class="p">,</span>
    <span class="n">temperature_c</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span>
    <span class="n">charge_efficiency</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span>
    <span class="n">discharge_efficiency</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span>
<span class="p">):</span>
    <span class="s">"""
    Adjusts the charge and discharge efficiency of the battery based on the temperature.

    Args:
        temperature_c (float): The temperature in degrees Celsius.
        charge_efficiency (float): The current charge efficiency of the battery.
        discharge_efficiency (float): The current discharge efficiency of the battery.

    Returns:
        Tuple[float, float]: A tuple containing the adjusted charge efficiency and discharge efficiency.
    """</span>
    <span class="n">temp_effect</span> <span class="o">=</span> <span class="nb">abs</span><span class="p">(</span><span class="n">temperature_c</span> <span class="o">-</span> <span class="mi">25</span><span class="p">)</span> <span class="o">*</span> <span class="mf">0.01</span>
    <span class="n">new_charge_efficiency</span> <span class="o">=</span> <span class="nb">max</span><span class="p">(</span><span class="mf">0.5</span><span class="p">,</span> <span class="nb">min</span><span class="p">(</span><span class="n">charge_efficiency</span> <span class="o">-</span> <span class="n">temp_effect</span><span class="p">,</span> <span class="mf">1.0</span><span class="p">))</span>
    <span class="n">new_discharge_efficiency</span> <span class="o">=</span> <span class="nb">max</span><span class="p">(</span>
        <span class="mf">0.5</span><span class="p">,</span> <span class="nb">min</span><span class="p">(</span><span class="n">discharge_efficiency</span> <span class="o">-</span> <span class="n">temp_effect</span><span class="p">,</span> <span class="mf">1.0</span><span class="p">)</span>
    <span class="p">)</span>
    <span class="k">return</span> <span class="n">new_charge_efficiency</span><span class="p">,</span> <span class="n">new_discharge_efficiency</span>
</code></pre></div></div>
<p>In the accompanying graph, observe the progression of State of Charge (SOC) and State of Health (SOH) values over time for each temperature, illustrating the dynamic behavior of the battery system.
<img src="https://raw.githubusercontent.com/koulakhilesh/OpenEnergy/master/images/notebook/assets/battery_simulation_temperature.png" alt="SOC and SOH" /></p>

<h2 id="battery">Battery</h2>

<p>This is the main class that represents a battery. It has various properties like capacity, efficiency, state of charge (SOC), state of health (SOH), and temperature. It also has methods for charging and discharging the battery, which adjust the efficiencies and update the SOH and cycle count.</p>

<p>The <code class="language-plaintext highlighter-rouge">charge</code> and <code class="language-plaintext highlighter-rouge">discharge</code> methods first adjust the efficiencies based on the temperature, then calculate the actual energy that can be charged or discharged, and finally update the SOC. They also call the <code class="language-plaintext highlighter-rouge">update_soh_and_cycles</code> method to update the SOH and cycle count.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">charge</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">energy_mwh</span><span class="p">:</span> <span class="nb">float</span><span class="p">):</span>
    <span class="s">"""
    Charges the battery with the specified amount of energy.

    Args:
        energy_mwh (float): The amount of energy to be charged in megawatt-hours.
    """</span>
    <span class="bp">self</span><span class="p">.</span><span class="n">adjust_efficiency_for_temperature</span><span class="p">()</span>
    <span class="n">energy_mwh</span> <span class="o">=</span> <span class="nb">min</span><span class="p">(</span><span class="n">energy_mwh</span><span class="p">,</span> <span class="bp">self</span><span class="p">.</span><span class="n">max_charge_rate_mw</span> <span class="o">*</span> <span class="bp">self</span><span class="p">.</span><span class="n">duration_hours</span><span class="p">)</span>
    <span class="n">actual_energy_mwh</span> <span class="o">=</span> <span class="n">energy_mwh</span> <span class="o">*</span> <span class="bp">self</span><span class="p">.</span><span class="n">charge_efficiency</span>
    <span class="bp">self</span><span class="p">.</span><span class="n">soc</span> <span class="o">=</span> <span class="nb">min</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">soc</span> <span class="o">+</span> <span class="n">actual_energy_mwh</span> <span class="o">/</span> <span class="bp">self</span><span class="p">.</span><span class="n">capacity_mwh</span><span class="p">,</span> <span class="mf">1.0</span><span class="p">)</span>
    <span class="bp">self</span><span class="p">.</span><span class="n">update_soh_and_cycles</span><span class="p">(</span><span class="n">energy_mwh</span><span class="p">)</span>
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">update_soh_and_cycles</code> method updates the SOH using the <code class="language-plaintext highlighter-rouge">BasicSOHCalculator</code> and updates the cycle count based on the energy cycled.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">update_soh_and_cycles</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">energy_mwh</span><span class="p">:</span> <span class="nb">float</span><span class="p">):</span>
    <span class="s">"""
    Updates the state of health (SOH) and cycle count of the battery based on the energy cycled.

    Args:
        energy_mwh (float): The amount of energy cycled in megawatt-hours.
    """</span>    
    <span class="bp">self</span><span class="p">.</span><span class="n">energy_cycled_mwh</span> <span class="o">+=</span> <span class="n">energy_mwh</span>
    <span class="n">dod</span> <span class="o">=</span> <span class="mf">1.0</span> <span class="o">-</span> <span class="bp">self</span><span class="p">.</span><span class="n">soc</span>
    <span class="bp">self</span><span class="p">.</span><span class="n">soh</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">soh_calculator</span><span class="p">.</span><span class="n">calculate_soh</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">soh</span><span class="p">,</span> <span class="n">energy_mwh</span><span class="p">,</span> <span class="n">dod</span><span class="p">)</span>
    <span class="bp">self</span><span class="p">.</span><span class="n">check_and_update_cycles</span><span class="p">()</span>
</code></pre></div></div>

<p>To explore the battery module in action, check out the Jupyter notebook <a href="https://github.com/koulakhilesh/OpenEnergy/blob/master/notebooks/assets/battery.ipynb">here</a>, where its functionality is demonstrated.</p>

<h2 id="wrapping-up">Wrapping Up</h2>

<p>That’s a brief overview of the <code class="language-plaintext highlighter-rouge">battery.py</code> module in our OpenEnergy project. We’ve covered how the module models a battery’s behavior, including how it adjusts efficiency based on temperature, calculates the state of health, and handles charging and discharging operations. I hope this gives you a better understanding of how we’re simulating energy storage systems in our project. Stay tuned for more deep dives into other parts of the codebase!</p>

<p>As always, if you have any questions or suggestions, feel free to leave a comment below or open an issue on our <a href="https://github.com/koulakhilesh/OpenEnergy/">GitHub repository</a>. Happy coding!</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="OpenEnergy" /><category term="Battery Module" /><category term="Python" /><category term="Energy Storage" /><category term="Simulation" /><category term="Efficiency Adjustment" /><category term="State of Health" /><category term="Charging" /><category term="Discharging" /><summary type="html"><![CDATA[Hello everyone! Today, we’re going to dive into the battery.py module of our OpenEnergy project. This module is the heart of our energy storage system simulation, and it’s where all the magic happens. You can find the complete code in our GitHub repository. This module is a great example of how to model a battery’s behavior in Python.]]></summary></entry><entry><title type="html">Understanding the Price Generation in OpenEnergy</title><link href="https://koulakhilesh.github.io/writing/understanding-the-price-generation/" rel="alternate" type="text/html" title="Understanding the Price Generation in OpenEnergy" /><published>2024-06-04T00:00:00+01:00</published><updated>2024-06-04T00:00:00+01:00</updated><id>https://koulakhilesh.github.io/writing/understanding-the-price-generation</id><content type="html" xml:base="https://koulakhilesh.github.io/writing/understanding-the-price-generation/"><![CDATA[<p>Hello everyone! In this post, we’ll dive into the price generation in the OpenEnergy project. We’ll be focusing on the <code class="language-plaintext highlighter-rouge">prices</code> folder under the <code class="language-plaintext highlighter-rouge">scripts</code> directory. You can find the code <a href="https://github.com/koulakhilesh/OpenEnergy/">here</a>.</p>

<h2 id="overview">Overview</h2>

<p>This module combines the functionality of various pricing strategies including simulated price, average price, and forecasted price. It provides a unified interface for interacting with different pricing models.</p>

<h2 id="interface">Interface</h2>

<p>The interface serves as a contract for all pricing models. It defines the methods that all pricing models should implement. This ensures that regardless of the pricing model used, the interaction remains consistent.</p>

<p>The <code class="language-plaintext highlighter-rouge">interfaces.py</code> file contains five interfaces: <code class="language-plaintext highlighter-rouge">IPriceData</code>, <code class="language-plaintext highlighter-rouge">IPriceEnvelopeGenerator</code>, <code class="language-plaintext highlighter-rouge">IPriceNoiseAdder</code>, and <code class="language-plaintext highlighter-rouge">IPriceDataHelper</code>. Each interface defines a set of methods that classes implementing the interface must provide.</p>

<h3 id="ipricedata">IPriceData</h3>

<p>The <code class="language-plaintext highlighter-rouge">IPriceData</code> interface is used for classes that retrieve price data. It defines a single method, <code class="language-plaintext highlighter-rouge">get_prices</code>, which takes a date as input and returns a tuple containing two lists of floats representing the buy and sell prices for that date.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">IPriceData</span><span class="p">(</span><span class="n">ABC</span><span class="p">):</span>
    <span class="s">"""Interface for retrieving price data."""</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_prices</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">Tuple</span><span class="p">[</span><span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">],</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]]:</span>
        <span class="s">"""Get the prices for a specific date.

        Args:
            date (datetime.date): The date for which to retrieve the prices.

        Returns:
            Tuple[List[float], List[float]]: A tuple containing two lists of floats.
                The first list represents the buy prices, and the second list represents
                the sell prices.
        """</span>
        <span class="k">pass</span>
</code></pre></div></div>

<h3 id="ipriceenvelopegenerator">IPriceEnvelopeGenerator</h3>

<p>The <code class="language-plaintext highlighter-rouge">IPriceEnvelopeGenerator</code> interface is used for classes that generate price envelopes. It defines a single method, <code class="language-plaintext highlighter-rouge">generate</code>, which takes a date as input and returns a list of price envelopes for that date.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">IPriceEnvelopeGenerator</span><span class="p">(</span><span class="n">ABC</span><span class="p">):</span>
    <span class="s">"""
    Interface for generating price envelopes.
    """</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">generate</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
        <span class="s">"""
        Generate price envelopes for the given date.

        Args:
            date (datetime.date): The date for which to generate price envelopes.

        Returns:
            List[float]: A list of price envelopes.
        """</span>
        <span class="k">pass</span>
</code></pre></div></div>

<h3 id="ipricenoiseadder">IPriceNoiseAdder</h3>

<p>The <code class="language-plaintext highlighter-rouge">IPriceNoiseAdder</code> interface is used for classes that add noise to a list of prices. It defines a single method, <code class="language-plaintext highlighter-rouge">add</code>, which takes a list of prices as input and returns a new list of prices with added noise.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">IPriceNoiseAdder</span><span class="p">(</span><span class="n">ABC</span><span class="p">):</span>
    <span class="s">"""
    Interface for adding noise to a list of prices.
    """</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">add</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">prices</span><span class="p">:</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">])</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
        <span class="s">"""
        Adds noise to the given list of prices.

        Args:
            prices (List[float]): The list of prices to add noise to.

        Returns:
            List[float]: The list of prices with added noise.
        """</span>
        <span class="k">pass</span>
</code></pre></div></div>

<h3 id="ipricedatahelper">IPriceDataHelper</h3>

<p>The [<code class="language-plaintext highlighter-rouge">IPriceDataHelper</code>] interface is designed for classes that assist in handling price data, offering a suite of methods for date and price data manipulation. This interface facilitates obtaining the current date and time, calculating the date and time a week prior to a specified date, retrieving data for the preceding week, fetching data for the current date, and extracting prices for the current date.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">IPriceDataHelper</span><span class="p">(</span><span class="n">ABC</span><span class="p">):</span>
    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_current_date</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">:</span>
        <span class="s">"""Retrieve the current date and time based on a given date.

        Args:
            date (datetime.date): The date to convert to datetime.

        Returns:
            datetime.datetime: The current date and time.
        """</span>
        <span class="k">pass</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_prior_date</span><span class="p">(</span>
        <span class="bp">self</span><span class="p">,</span> <span class="n">current_date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">,</span> <span class="n">delta_days</span><span class="p">:</span> <span class="nb">int</span>
    <span class="p">)</span> <span class="o">-&gt;</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">:</span>
        <span class="s">"""Calculate the date prior to the current date by a specified number of days.

        Args:
            current_date (datetime.datetime): The current date from which to calculate the prior date.
            delta_days (int): The number of days before the current date to calculate.

        Returns:
            datetime.datetime: The calculated prior date.
        """</span>
        <span class="k">pass</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_prior_data</span><span class="p">(</span>
        <span class="bp">self</span><span class="p">,</span>
        <span class="n">current_date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">,</span>
        <span class="n">prior_date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">,</span>
        <span class="n">data</span><span class="p">:</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span><span class="p">,</span>
    <span class="p">)</span> <span class="o">-&gt;</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span><span class="p">:</span>
        <span class="s">"""Retrieve data for a period between the prior date and the current date from a given dataset.

        Args:
            current_date (datetime.datetime): The end date of the period for which to retrieve data.
            prior_date (datetime.datetime): The start date of the period for which to retrieve data.
            data (pd.DataFrame): The dataset from which to retrieve the data.

        Returns:
            pd.DataFrame: The filtered dataset containing data between the prior and current dates.
        """</span>
        <span class="k">pass</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_current_date_data</span><span class="p">(</span>
        <span class="bp">self</span><span class="p">,</span> <span class="n">current_date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">datetime</span><span class="p">,</span> <span class="n">data</span><span class="p">:</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span>
    <span class="p">)</span> <span class="o">-&gt;</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span><span class="p">:</span>
        <span class="s">"""Retrieve data for the current date from a given dataset.

        Args:
            current_date (datetime.datetime): The date for which to retrieve data.
            data (pd.DataFrame): The dataset from which to retrieve the data.

        Returns:
            pd.DataFrame: The filtered dataset containing data for the current date.
        """</span>
        <span class="k">pass</span>

    <span class="o">@</span><span class="n">abstractmethod</span>
    <span class="k">def</span> <span class="nf">get_prices_current_date</span><span class="p">(</span>
        <span class="bp">self</span><span class="p">,</span> <span class="n">current_date_data</span><span class="p">:</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span><span class="p">,</span> <span class="n">column_name</span><span class="p">:</span> <span class="nb">str</span>
    <span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
        <span class="s">"""Extract price data for the current date from a given dataset.

        Args:
            current_date_data (pd.DataFrame): The dataset containing data for the current date.
            column_name (str): The name of the column from which to extract the prices.

        Returns:
            List[float]: The list of prices for the current date.
        """</span>
        <span class="k">pass</span>
</code></pre></div></div>

<h2 id="simulated-price">Simulated Price</h2>

<p>The simulated price model uses statistical methods to generate a price that simulates market conditions. It’s useful for testing how systems would react to different market conditions.</p>

<p>The file we’ll look at is <code class="language-plaintext highlighter-rouge">simulated_price.py</code>. This file contains the logic for generating simulated price data. It’s a crucial part of the project as it allows us to create realistic, yet artificial, price data for testing and development purposes.</p>

<h3 id="simulatedpriceenvelopegenerator">SimulatedPriceEnvelopeGenerator</h3>

<p>The <code class="language-plaintext highlighter-rouge">SimulatedPriceEnvelopeGenerator</code> class is responsible for generating a simulated price envelope based on a sine wave. The envelope is defined by a number of intervals, a minimum and maximum price, and a peak start and end index.</p>

<p>The <code class="language-plaintext highlighter-rouge">generate</code> method is where the magic happens. It generates a list of price values for a given date. The prices are calculated based on the position of the interval in relation to the peak start and end. If the interval is within the peak, the price is calculated using a sine wave function. If it’s outside the peak, the price is calculated using a smaller amplitude sine wave. A random adjustment is also added to each price to simulate real-world price fluctuations.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">SimulatedPriceEnvelopeGenerator</span><span class="p">(</span><span class="n">IPriceEnvelopeGenerator</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">generate</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
        <span class="s">"""
        Generates a simulated price envelope for the given date.

        Args:
            date (datetime.date): The date for which to generate the price envelope.

        Returns:
            List[float]: A list of price values representing the price envelope.
        """</span>
        <span class="c1"># ... code omitted for brevity ...
</span>        <span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">num_intervals</span><span class="p">):</span>
            <span class="c1"># ... code omitted for brevity ...
</span>            <span class="n">prices</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">price</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">prices</span>
</code></pre></div></div>

<h3 id="simulatedpricenoiseadder">SimulatedPriceNoiseAdder</h3>

<p>The <code class="language-plaintext highlighter-rouge">SimulatedPriceNoiseAdder</code> class adds simulated noise to the list of prices generated by the <code class="language-plaintext highlighter-rouge">SimulatedPriceEnvelopeGenerator</code>. This is done to make the simulated prices more realistic. The noise is added by randomly adjusting each price within a specified noise level. There’s also a chance for a price spike to occur, which multiplies the price by a specified spike multiplier.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">SimulatedPriceNoiseAdder</span><span class="p">(</span><span class="n">IPriceNoiseAdder</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">add</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">prices</span><span class="p">:</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">])</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
        <span class="s">"""
        Adds simulated noise to a list of prices.

        Args:
            prices (List[float]): The list of prices to add noise to.

        Returns:
            List[float]: The list of prices with simulated noise added.
        """</span>
        <span class="c1"># ... code omitted for brevity ...
</span>        <span class="k">for</span> <span class="n">price</span> <span class="ow">in</span> <span class="n">prices</span><span class="p">:</span>
            <span class="c1"># ... code omitted for brevity ...
</span>            <span class="n">noisy_prices</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">new_price</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">noisy_prices</span>
</code></pre></div></div>

<h3 id="simulatedpricemodel">SimulatedPriceModel</h3>

<p>The <code class="language-plaintext highlighter-rouge">SimulatedPriceModel</code> class brings it all together. It uses an instance of <code class="language-plaintext highlighter-rouge">SimulatedPriceEnvelopeGenerator</code> to generate a list of prices for a given date, and then adds noise to these prices using an instance of <code class="language-plaintext highlighter-rouge">SimulatedPriceNoiseAdder</code>. The result is a tuple containing two lists of prices: one with the original prices and one with the noisy prices.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">SimulatedPriceModel</span><span class="p">(</span><span class="n">IPriceData</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">get_prices</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">Tuple</span><span class="p">[</span><span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">],</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]]:</span>
        <span class="s">"""
        Generates simulated prices for the given date.

        Args:
            date (datetime.date): The date for which prices need to be generated.

        Returns:
            Tuple[List[float], List[float]]: A tuple containing two lists of prices.
            The first list represents the prices without noise and spikes,
            and the second list represents the prices with noise and spikes.

        """</span>
        <span class="c1"># ... code omitted for brevity ...
</span>        <span class="n">prices_with_noise_and_spikes</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">noise_adder</span><span class="p">.</span><span class="n">add</span><span class="p">(</span><span class="n">prices</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">prices</span><span class="p">,</span> <span class="n">prices_with_noise_and_spikes</span>
</code></pre></div></div>

<p>Below is the graph depicting the outcomes from the SimulatedPriceModel, showcasing the simulated prices:
<img src="https://raw.githubusercontent.com/koulakhilesh/OpenEnergy/master/images/notebook/prices/simulated_prices.png" alt="Simulated Prices" /></p>

<p>To explore the more on SimulatedPriceModel , check out the Jupyter notebook <a href="https://github.com/koulakhilesh/OpenEnergy/blob/master/notebooks/prices/simulated_price.ipynb">here</a>, where its functionality is demonstrated.</p>

<h2 id="average-price">Average Price</h2>

<p>The average price model calculates the average price based on historical data. It’s a simple yet effective model for predicting future prices when the market conditions are stable.</p>

<p>The <code class="language-plaintext highlighter-rouge">average_price.py</code> file contains the <code class="language-plaintext highlighter-rouge">HistoricalAveragePriceModel</code> class, which calculates historical average prices. This class is a key component of the project as it provides the historical context needed to understand current price data.</p>

<h3 id="historicalaveragepricemodel">HistoricalAveragePriceModel</h3>

<p>The <code class="language-plaintext highlighter-rouge">HistoricalAveragePriceModel</code> class is initialized with a data provider and an optional boolean indicating whether to interpolate missing values in the data. The data provider is used to retrieve price data, and the interpolation option determines whether missing values in the data are filled in using linear interpolation.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">HistoricalAveragePriceModel</span><span class="p">(</span><span class="n">IPriceData</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">data_provider</span><span class="p">:</span> <span class="n">IDataProvider</span><span class="p">,</span> <span class="n">interpolate</span><span class="p">:</span> <span class="nb">bool</span> <span class="o">=</span> <span class="bp">True</span><span class="p">,</span> <span class="n">prior_days</span><span class="p">:</span> <span class="nb">int</span> <span class="o">=</span> <span class="n">DAYS_IN_WEEK</span><span class="p">):</span>
        <span class="c1"># ... code omitted for brevity ...
</span>        <span class="k">if</span> <span class="bp">self</span><span class="p">.</span><span class="n">interpolate</span><span class="p">:</span>
            <span class="bp">self</span><span class="p">.</span><span class="n">data</span><span class="p">[</span><span class="bp">self</span><span class="p">.</span><span class="n">PRICE_COLUMN</span><span class="p">].</span><span class="n">interpolate</span><span class="p">(</span><span class="n">method</span><span class="o">=</span><span class="s">"linear"</span><span class="p">,</span> <span class="n">inplace</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
</code></pre></div></div>

<h3 id="get_prices">get_prices</h3>

<p>The <code class="language-plaintext highlighter-rouge">get_prices</code> method is the main method of the class. It takes a date as input and returns a tuple containing the average prices for the last week and the prices for the current date. The method uses helper methods to get the current date, the date a week prior, the data for the last week, the average prices for the last week, the data for the current date, and the prices for the current date.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">get_prices</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">Tuple</span><span class="p">[</span><span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">],</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]]:</span>
    <span class="s">"""
    Get the average prices for the last week and the prices for the current date.

    Args:
        date (datetime.date): The current date.

    Returns:
        Tuple[List[float], List[float]]: A tuple containing the average prices for the last week
        and the prices for the current date.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>    <span class="k">return</span> <span class="n">average_prices_last_week</span><span class="p">,</span> <span class="n">prices_current_date</span>
</code></pre></div></div>

<h3 id="get_average_prices_last_week">get_average_prices_last_week</h3>

<p>The <code class="language-plaintext highlighter-rouge">get_average_prices_last_week</code> method calculates the average prices for the last week. It takes a DataFrame containing the price data for the last week as input and returns a list of average prices for each hour of the day. The method groups the data by hour and calculates the mean price for each group.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">get_average_prices_last_week</span><span class="p">(</span>
    <span class="bp">self</span><span class="p">,</span> <span class="n">last_week_data</span><span class="p">:</span> <span class="n">pd</span><span class="p">.</span><span class="n">DataFrame</span>
<span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]:</span>
    <span class="s">"""
    Calculate the average prices for the last week.

    Args:
        last_week_data (pd.DataFrame): The price data for the last week.

    Returns:
        List[float]: A list of average prices for each hour of the day.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>    <span class="k">return</span> <span class="p">(</span>
        <span class="n">last_week_data</span><span class="p">.</span><span class="n">groupby</span><span class="p">(</span><span class="n">last_week_data</span><span class="p">.</span><span class="n">index</span><span class="p">.</span><span class="n">hour</span><span class="p">)[</span><span class="bp">self</span><span class="p">.</span><span class="n">PRICE_COLUMN</span><span class="p">]</span>
        <span class="p">.</span><span class="n">mean</span><span class="p">()</span>
        <span class="p">.</span><span class="n">tolist</span><span class="p">()</span>
    <span class="p">)</span>
</code></pre></div></div>

<p>Below is the graph depicting the mean and standard deviation observed from the HistoricalAveragePriceModel, showcasing the simulated prices:
<img src="https://raw.githubusercontent.com/koulakhilesh/OpenEnergy/master/images/notebook/prices/Mean_and_Standard_Deviation_of_GB_GBN_price_day_ahead_by_Hour.png" alt="Mean and std" /></p>

<p>To explore the more on HistoricalAveragePriceModel , check out the Jupyter notebook <a href="https://github.com/koulakhilesh/OpenEnergy/blob/master/notebooks/prices/average_price.ipynb">here</a>, where its functionality is demonstrated.</p>

<h2 id="forecasted-price">Forecasted Price</h2>

<p>The forecasted price model uses advanced statistical methods or machine learning algorithms to predict future prices. It’s the most complex model and can adapt to changing market conditions.</p>

<p>The <code class="language-plaintext highlighter-rouge">forecasted_price.py</code> file contains the <code class="language-plaintext highlighter-rouge">ForecastPriceModel</code> class, which implements the <code class="language-plaintext highlighter-rouge">IPriceData</code> and <code class="language-plaintext highlighter-rouge">IForecaster</code> interfaces. This class provides methods for training, forecasting, and evaluating prices.</p>

<h3 id="forecastpricemodel">ForecastPriceModel</h3>

<p>The <code class="language-plaintext highlighter-rouge">ForecastPriceModel</code> class is initialized with a data provider, a feature engineer, a machine learning model, a history length, a forecast length, and an optional boolean indicating whether to interpolate missing values in the data. The data provider is used to retrieve price data, the feature engineer is used to perform feature engineering on the data, and the machine learning model is used for forecasting.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">ForecastPriceModel</span><span class="p">(</span><span class="n">IPriceData</span><span class="p">,</span> <span class="n">IForecaster</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span>
        <span class="bp">self</span><span class="p">,</span>
        <span class="n">data_provider</span><span class="p">:</span> <span class="n">IDataProvider</span><span class="p">,</span>
        <span class="n">feature_engineer</span><span class="p">:</span> <span class="n">IFeatureEngineer</span><span class="p">,</span>
        <span class="n">model</span><span class="p">:</span> <span class="n">IModel</span><span class="p">,</span>
        <span class="n">history_length</span><span class="o">=</span><span class="mi">7</span> <span class="o">*</span> <span class="mi">24</span><span class="p">,</span>
        <span class="n">forecast_length</span><span class="o">=</span><span class="mi">24</span><span class="p">,</span>
        <span class="n">interpolate</span><span class="p">:</span> <span class="nb">bool</span> <span class="o">=</span> <span class="bp">True</span><span class="p">,</span>
        <span class="n">prior_days</span><span class="p">:</span> <span class="nb">int</span> <span class="o">=</span> <span class="n">DAYS_IN_WEEK</span><span class="p">,</span>
    <span class="p">):</span>
        <span class="c1"># ... code omitted for brevity ...
</span></code></pre></div></div>

<h3 id="get_prices-1">get_prices</h3>

<p>The <code class="language-plaintext highlighter-rouge">get_prices</code> method is the main method of the class. It takes a date as input and returns a tuple containing the forecasted prices and the actual prices for the given date. The method uses helper methods to get the current date, the date a week prior, the data for the last week, the forecasted prices for the current date, the data for the current date, and the prices for the current date.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">get_prices</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">date</span><span class="p">:</span> <span class="n">datetime</span><span class="p">.</span><span class="n">date</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">t</span><span class="p">.</span><span class="n">Tuple</span><span class="p">[</span><span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">],</span> <span class="n">t</span><span class="p">.</span><span class="n">List</span><span class="p">[</span><span class="nb">float</span><span class="p">]]:</span>
    <span class="s">"""
    Get the forecasted prices and actual prices for a given date.

    Args:
        date (datetime.date): The date for which to get the prices.

    Returns:
        Tuple[List[float], List[float]]: A tuple containing the forecasted prices and actual prices.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>    <span class="k">return</span> <span class="n">forecasted_prices</span><span class="p">,</span> <span class="n">prices_current_date</span>
</code></pre></div></div>

<h3 id="train-forecast-evaluate-save_model-load_model">train, forecast, evaluate, save_model, load_model</h3>

<p>The <code class="language-plaintext highlighter-rouge">train</code>, <code class="language-plaintext highlighter-rouge">forecast</code>, <code class="language-plaintext highlighter-rouge">evaluate</code>, <code class="language-plaintext highlighter-rouge">save_model</code>, and <code class="language-plaintext highlighter-rouge">load_model</code> methods are used to train the machine learning model, forecast prices, evaluate the forecasted prices, save the trained model to a file, and load a trained model from a file, respectively.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">train</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">df</span><span class="p">):</span>
    <span class="s">"""
    Train the forecast price model.

    Args:
        df (pandas.DataFrame): The training data.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>
<span class="k">def</span> <span class="nf">forecast</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">df</span><span class="p">):</span>
    <span class="s">"""
    Forecast prices using the trained model.

    Args:
        df (pandas.DataFrame): The data to forecast.

    Returns:
        pandas.DataFrame: The forecasted prices.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>
<span class="k">def</span> <span class="nf">evaluate</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">y_true</span><span class="p">,</span> <span class="n">y_pred</span><span class="p">):</span>
    <span class="s">"""
    Evaluate the forecasted prices.

    Args:
        y_true (numpy.ndarray): The true prices.
        y_pred (numpy.ndarray): The forecasted prices.

    Returns:
        float: The evaluation metric.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>
<span class="k">def</span> <span class="nf">save_model</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">file_path</span><span class="p">):</span>
    <span class="s">"""
    Save the trained model to a file.

    Args:
        file_path (str): The path to the file.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span>
<span class="o">@</span><span class="nb">staticmethod</span>
<span class="k">def</span> <span class="nf">load_model</span><span class="p">(</span><span class="n">file_path</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">IModel</span><span class="p">:</span>
    <span class="s">"""
    Load a trained model from a file.

    Args:
        file_path (str): The path to the file.

    Returns:
        IModel: The loaded model.
    """</span>
    <span class="c1"># ... code omitted for brevity ...
</span></code></pre></div></div>

<p>Below is the graph depicting the outcomes from the ForecastPriceModel, showcasing the simulated prices:
<img src="https://raw.githubusercontent.com/koulakhilesh/OpenEnergy/master/images/notebook/prices/forecast_vs_actual.png" alt="Forecasted Prices" /></p>

<p>To explore the more on ForecastPriceModel , check out the Jupyter notebook <a href="https://github.com/koulakhilesh/OpenEnergy/blob/master/notebooks/prices/forecasted_price.ipynb">here</a>, where its functionality is demonstrated.</p>

<p>To use any of the pricing models, you need to create an instance of the model and call the appropriate methods as defined in the interface. The specific implementation details depend on the programming language and the design of your software.</p>

<p>Please note that this is a high-level overview. For detailed information, refer to the specific documentation for each pricing model and the source code.</p>

<p>Remember, the best way to learn is by doing. So, I encourage you to clone the <a href="https://github.com/koulakhilesh/OpenEnergy/">repository</a>, play around with the code, and see what you can create. Happy coding!</p>]]></content><author><name>Akhilesh Koul</name><email>koulakhilesh@gmail.com</email></author><category term="OpenEnergy" /><category term="Price Generation" /><category term="Python" /><category term="Pricing Models" /><category term="Simulated Price" /><category term="Average Price" /><category term="Forecasted Price" /><category term="Machine Learning" /><summary type="html"><![CDATA[Hello everyone! In this post, we’ll dive into the price generation in the OpenEnergy project. We’ll be focusing on the prices folder under the scripts directory. You can find the code here.]]></summary></entry></feed>