<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Technology &#8211; Thejesh GN</title>
	<atom:link href="https://thejeshgn.com/category/technology/feed/" rel="self" type="application/rss+xml" />
	<link>https://thejeshgn.com</link>
	<description>A container for all my views with excerpts from technology, travel, films, books, kannada, friends and other interests. I am Thejesh GN, friends call me Thej.</description>
	<lastBuildDate>Tue, 14 Jul 2026 07:47:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.9.5</generator>

<image>
	<url>https://thejeshgn.com/wp-content/uploads/2015/08/cropped-thejeshgn_icon1-150x150.png</url>
	<title>Technology &#8211; Thejesh GN</title>
	<link>https://thejeshgn.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">9645742</site>	<item>
		<title>Linked List : Annotating Photos for Humans and Machines</title>
		<link>https://thejeshgn.com/2026/07/14/linked-list-annotating-photos-for-humans-and-machines/</link>
					<comments>https://thejeshgn.com/2026/07/14/linked-list-annotating-photos-for-humans-and-machines/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Tue, 14 Jul 2026 07:47:25 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[Free and Open Source]]></category>
		<category><![CDATA[Linked List]]></category>
		<category><![CDATA[Open Data]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39312</guid>

					<description><![CDATA[Photos taken on a mobile phone can include metadata. I like it for various reasons. My various projects depend on metadata, such as date, location, orientation, etc. But more and more, the phones are removing that metadata when I transmit it. Some of them also change the file name. But I care, and&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>Photos taken on a mobile phone can include metadata. I like it for various reasons. My various projects depend on metadata, such as date, location, orientation, etc. But more and more, the phones are removing that metadata when I transmit it. Some of them also change the file name. But I care, and I want them. It&#8217;s a carefully considered call.</p>



<p>In response to these aggressive moves by the platforms, I have started adding the data directly to the visible part of the image, like how your old film camera would add the date/time directly onto the photo itself. There is no danger of losing that information. As long as I have the picture, I have the date/time it was taken.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><a href="https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537.jpg" rel="lightbox[39312]"><img fetchpriority="high" decoding="async" width="2000" height="1125" src="https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537.jpg" alt="Gummalapuram Branch P.O" class="wp-image-39104" srcset="https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537.jpg 2000w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-300x169.jpg 300w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-1024x576.jpg 1024w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-768x432.jpg 768w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-1536x864.jpg 1536w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-720x405.jpg 720w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-520x293.jpg 520w, https://thejeshgn.com/wp-content/uploads/2026/06/wp-17820131989467051048442338121537-320x180.jpg 320w" sizes="(max-width: 2000px) 100vw, 2000px" /></a><figcaption class="wp-element-caption">Gummalapuram Branch P.O</figcaption></figure>
</div>


<p>Similarly, now, when I take certain pictures, I add text directly to the bottom-right corner, usually in English. I am making sure it has good contrast so humans and OCRs can pick it up easily. I have scripts to extract this data easily if I want to process it. Of course, any cheap visual model can also do it, but it&#8217;s easy and cheap to make the model write code and reuse it.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><a href="https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577.png" rel="lightbox[39312]"><img decoding="async" width="900" height="2000" src="https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577.png" alt="6oz Coffee" class="wp-image-39309" style="width:auto;height:500px" srcset="https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577.png 900w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-135x300.png 135w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-461x1024.png 461w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-768x1707.png 768w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-691x1536.png 691w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-720x1600.png 720w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-520x1156.png 520w, https://thejeshgn.com/wp-content/uploads/2026/07/wp-17839700613736043349321074350577-320x711.png 320w" sizes="(max-width: 900px) 100vw, 900px" /></a><figcaption class="wp-element-caption">6oz Coffee</figcaption></figure>
</div>


<p>I have also started playing around with adding text to a QR code and then placing it on top of a Photo/Image. It works very well. It makes it much easier for machines to read and can also be read by humans using almost any phone camera.</p>



<p>The Apps that I use are</p>



<ol class="wp-block-list">
<li><a href="https://opencamera.org.uk/" target="_blank" rel="noreferrer noopener">OpenCamera</a> (<a href="https://f-droid.org/en/packages/net.sourceforge.opencamera/" target="_blank" rel="noreferrer noopener">F-Droid</a>, <a href="https://play.google.com/store/apps/details?id=net.sourceforge.opencamera&amp;hl=en_IN" target="_blank" rel="noreferrer noopener">Play</a>): The most flexible camera app I have used. It&#8217;s open source. I use the GPS and Text stamping features a lot. It adds GPS tags, a timestamp, and a line of text on top of the image. I wish it had some more text features, but it works. All my new <a href="https://thejeshgn.com/projects/openpostboxindia/" target="_blank" rel="noreferrer noopener">OpenPostboxIndia</a> project pictures are taken with this. So unlike my first set of pictures, however they are maintained or transmitted, that data will remain with them. It can also be configured to take a sequence of pictures (like a timelapse) that can be uploaded to <a href="https://panoramax.fr/comment-participer-a-panoramax" target="_blank" rel="noreferrer noopener">Panoramax</a>. I also love that it can record video along with the GPS tags as a separate subtitle text.</li>



<li><a href="https://github.com/T8RIN/ImageToolbox" target="_blank" rel="noreferrer noopener">ImageToolBox</a> (<a href="https://f-droid.org/packages/ru.tech.imageresizershrinker/" target="_blank" rel="noreferrer noopener">FDroid</a>, Obtanium and <a href="https://play.google.com/store/apps/details?id=ru.tech.imageresizershrinker" target="_blank" rel="noreferrer noopener">Play</a>): A FOSS Swiss Army knife of image editing and annotating images or pictures on Phone. It has a <a href="https://github.com/T8RIN/ImageToolbox#-features" target="_blank" rel="noreferrer noopener">ton of tools</a>; some aren&#8217;t even directly related to image manipulation, like PDF Tools. So take your time to check out all the tools. Recently I discovered OCR tools and some local AI-based tools. All useful. I generally use the built-in QR code generator, then the watermarking feature to add that QR code to the image. Sometimes I use it for hand drawn notes or markers on top of pictures. I am still exploring this app. I am sure I will fit it into my other workflows and uninstall some other apps.</li>



<li><a href="https://github.com/markusfisch/BinaryEye" target="_blank" rel="noreferrer noopener">BinaryEye</a> (<a href="https://f-droid.org/en/packages/de.markusfisch.android.binaryeye/" target="_blank" rel="noreferrer noopener">FDroid</a>, <a href="https://play.google.com/store/apps/details?id=de.markusfisch.android.binaryeye" target="_blank" rel="noreferrer noopener">Play</a>) is a dedicated QR Code (Barcode) scanner and creator App. It works well and can also do other things, like forwarding the scanned content to a a web endpoint or working as a remote Bluetooth scanner.</li>
</ol>



<p>Do you annotate photographs or pictures? If yes, why? and What tools do you use?</p>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39312" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/07/14/linked-list-annotating-photos-for-humans-and-machines/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">39312</post-id>	</item>
		<item>
		<title>SIR Online: Everything Is Difficult by Design</title>
		<link>https://thejeshgn.com/2026/07/09/sir-online-everything-is-difficult-by-design/</link>
					<comments>https://thejeshgn.com/2026/07/09/sir-online-everything-is-difficult-by-design/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Thu, 09 Jul 2026 17:52:18 +0000</pubDate>
				<category><![CDATA[Life]]></category>
		<category><![CDATA[Technology]]></category>
		<category><![CDATA[Aadhaar]]></category>
		<category><![CDATA[Elections]]></category>
		<category><![CDATA[Politics]]></category>
		<location><![CDATA[India 🇮🇳]]></location>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39277</guid>

					<description><![CDATA[I did SIR enumeration online. The application seems to have been written by someone who hates users and doesn&#8217;t want them to get the work done1. I have had the same EPIC number since 2002, and it is still valid. I have kept my information updated over the years. Yet, I had to&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>I did SIR enumeration online. The application seems to have been written by someone who hates users and doesn&#8217;t want them to get the work done<sup class='footnote'><a href='#fn-39277-1' id='fnref-39277-1' onclick='return fdfootnote_show(39277)'>1</a></sup>.</p>



<p>I have had the same EPIC number since 2002, and it is still valid. I have kept my information updated over the years. Yet, I had to search for myself in a 2002 electoral roll, and I couldn&#8217;t find my name. I don&#8217;t exactly remember what happened in 2002 or why my name isn&#8217;t there. Or may be it was added after 2002 SIR<sup class='footnote'><a href='#fn-39277-2' id='fnref-39277-2' onclick='return fdfootnote_show(39277)'>2</a></sup>. Ideally system should know it and guide accordingly.</p>



<p>Then I had to search for one of my parents and link their details. Thanks to my father&#8217;s good memory we were able to find the part no. Because the search dropdowns doesn&#8217;t always show the polling station address. One has to know it!</p>


<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search.png" rel="lightbox[39277]"><img decoding="async" width="1024" height="602" src="https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-1024x602.png" alt="Searching your name in SIR in 2002." class="wp-image-39278" srcset="https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-1024x602.png 1024w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-300x176.png 300w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-768x452.png 768w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-1536x903.png 1536w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-2048x1204.png 2048w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-720x423.png 720w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-520x306.png 520w, https://thejeshgn.com/wp-content/uploads/2012/09/sir_2002_search-320x188.png 320w" sizes="(max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Searching your name in SIR in 2002.</figcaption></figure>
</div>


<p>Of course, Aadhaar is not compulsory<sup class='footnote'><a href='#fn-39277-3' id='fnref-39277-3' onclick='return fdfootnote_show(39277)'>3</a></sup>; it&#8217;s just that you can&#8217;t submit the form online without e-signing it using Aadhaar + OTP. I have a valid digital signature certificate that I use for e-signing. Why can&#8217;t I use that? No one knows.</p>



<p>What happens after the submission, no one knows. There is acknowledgement shown on the screen after the signing, but I dont know what happens after that. You do get a SMS message of the acknowledgement. But that message also doesn&#8217;t show any acknowledgement or any details on how to check on it.</p>



<p>For someone with a tech background, it took considerable amount time and effort to submit the form. Imagine folks who are less comfortable with technology or those who have changed locations and doesn&#8217;t remember well. Imagine those who fall for things like “digital arrest” scams trying to navigate this.</p>



<p>What a mess this #SIR is. The #ECI website is probably one of the most user-unfriendly applications I have seen online recently. I thought IRCTC was bad.</p>



<p>Some coverage</p>



<ul class="wp-block-list">
<li><a href="https://citizenmatters.in/sir-for-karnataka-voters-all-you-need-to-know-about-enumeration/">Citizen Matter&#8217;s has an How &#8211; To</a></li>



<li><a href="https://www.thehindu.com/news/cities/bangalore/sir-voters-blos-unsure-what-follows-after-filling-online-form/article71183278.ece">The Hindu &#8211; SIR: Voters, BLOs unsure what follows after filling online form</a></li>



<li><a href="http://timesofindia.indiatimes.com/articleshow/132245659.cms?utm_source=thejeshgn.com">TOI &#8211; SIR’s digital wall: 1 lakh try in Karnataka, only 0.2% get through</a></li>



<li><a href="https://www.thehindu.com/news/cities/bangalore/sir-progeny-mapping-in-karnataka-different-blos-different-instructions/article71179256.ece">The Hindu &#8211; SIR progeny mapping in Karnataka: Different BLOs, different instructions</a></li>



<li><a href="https://www.thehindu.com/news/cities/bangalore/sir-missing-2002-supplementary-rolls-leave-voters-unable-to-verify-electoral-records/article71211009.ece">The Hindu &#8211; SIR: Missing 2002 ‘supplementary rolls’ leave voters unable to verify electoral records</a></li>
</ul>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39277" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>



<h3 class="wp-block-heading">Footnotes</h3>


<div class='footnotes' id='footnotes-39277'><div class='footnotedivider'></div><ol><li id='fn-39277-1'> Your mobile number has to associated with your current EPIC or you can&#8217;t. To do that you need to submit Form 8 <span class='footnotereverse'><a href='#fnref-39277-1'>&#8617;</a></span></li><li id='fn-39277-2'> Learnt later, as per the The Hindu,  2002 supplementary rolls are not included, so may be that is the reason? <span class='footnotereverse'><a href='#fnref-39277-2'>&#8617;</a></span></li><li id='fn-39277-3'> Also remember the names have to match exactly between your current EPIC and Aadhaar. If they don&#8217;t match then submit the enumeration form offline. <span class='footnotereverse'><a href='#fnref-39277-3'>&#8617;</a></span></li></ol></div>]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/07/09/sir-online-everything-is-difficult-by-design/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">39277</post-id>	</item>
		<item>
		<title>Getting Started with SimulIDE and Arduino</title>
		<link>https://thejeshgn.com/2026/07/03/getting-started-with-simulide-and-arduino/</link>
					<comments>https://thejeshgn.com/2026/07/03/getting-started-with-simulide-and-arduino/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Thu, 02 Jul 2026 20:02:14 +0000</pubDate>
				<category><![CDATA[Podcast]]></category>
		<category><![CDATA[Technology]]></category>
		<category><![CDATA[Basic Electronics]]></category>
		<category><![CDATA[Free and Open Source]]></category>
		<category><![CDATA[Hardware Simulation]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39215</guid>

					<description><![CDATA[I was looking for an alternative to Tinkercad to simulate building, coding, and running Arduino projects. SimulIDE Circuit Simulator seemed like an ideal candidate. SimulIDE is a simple real time electronic circuit simulator, intended for hobbyist or students to learn and experiment with analog and digital electronic circuits and microcontrollers. It supports PIC,&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>I was looking for an alternative to Tinkercad to simulate building, coding, and running Arduino projects. SimulIDE Circuit Simulator seemed like an ideal candidate.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p><a href="https://simulide.com/p/">SimulIDE</a> is a simple real time electronic circuit simulator, intended for hobbyist or students to learn and experiment with analog and digital electronic circuits and microcontrollers. It supports PIC, AVR , Arduino and other MCUs and MPUs.</p>
</blockquote>



<p>I installed it from Flathub (<a href="https://flathub.org/en/apps/com.simulide.simulide">SimulIDE</a>). It worked well. It is easy to understand and work with. It&#8217;s also easy for a beginner to pick up and start designing circuits in 5 minutes. I tried a few basic circuits. Then I tried the Blink circuit and Fade circuits based on Arduino. Unlike previous ones, this one included building a firmware and loading it. To make it complex, I was also using the Flatpak version of the Arduino IDE. Then I found two options</p>



<h3 class="wp-block-heading">Load Firmware in SimulIDE</h3>



<p>In this option, we will build the software in the Arduino IDE. Compile it. Once it generates a firmware (a HEX file), we will load it into SimulIDE. For this, we need to know where the compilation output is in the Arduino IDE, and then make sure SimulIDE has access to that path. Enable verbose output in the Arduino IDE Preferences, and then compile. The IDE will print the path. Take the path, go to <a href="https://flathub.org/en/apps/com.github.tchx84.Flatseal">Flatseal</a> <sup class='footnote'><a href='#fn-39215-1' id='fnref-39215-1' onclick='return fdfootnote_show(39215)'>1</a></sup>, and give SimulIDE access to that path.</p>


<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="635" src="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-1024x635.png" alt="Arduino IDE, Enable Verbose output" class="wp-image-39224" srcset="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-1024x635.png 1024w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-300x186.png 300w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-768x476.png 768w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-720x446.png 720w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-520x322.png 520w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose-320x198.png 320w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_enable_verbose.png 1511w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Arduino IDE, Enable Verbose output</figcaption></figure>
</div>

<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="650" src="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-1024x650.png" alt="Path of firmware hexfile built by Arduino IDE" class="wp-image-39225" srcset="https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-1024x650.png 1024w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-300x191.png 300w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-768x488.png 768w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-1536x975.png 1536w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-720x457.png 720w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-520x330.png 520w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub-320x203.png 320w, https://thejeshgn.com/wp-content/uploads/2026/07/arduino_ide_v2_hex_path_flathub.png 1951w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Path of firmware hexfile built by Arduino IDE</figcaption></figure>
</div>

<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="377" src="https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-1024x377.png" alt="Giving folder access to SimulIDE so it can access the Hex files in Flatseal." class="wp-image-39226" srcset="https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-1024x377.png 1024w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-300x110.png 300w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-768x283.png 768w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-1536x565.png 1536w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-720x265.png 720w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-520x191.png 520w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files-320x118.png 320w, https://thejeshgn.com/wp-content/uploads/2026/07/simulIDE_accessing_arduino_compiled_files.png 1892w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Giving folder access to SimulIDE so it can access the Hex files in Flatseal.</figcaption></figure>
</div>


<p>Now SimulIDE has access to the folders where firmware is built. Before you simulate the circuit, right-click the Uno, go to mega328-109, and then load firmware. Choose the hex file. Alternatively, right-click Uno, go to mega328-109 &#8211; Properties, enter the firmware path, and select Reload HEX at Simulation Start.</p>


<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="744" src="https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-1024x744.png" alt="Accessing Arduino Uno properties in SimulIDE" class="wp-image-39227" srcset="https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-1024x744.png 1024w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-300x218.png 300w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-768x558.png 768w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-720x523.png 720w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-520x378.png 520w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings-320x233.png 320w, https://thejeshgn.com/wp-content/uploads/2007/06/simulide_uno_settings.png 1285w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Accessing Arduino Uno properties in SimulIDE</figcaption></figure>
</div>

<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="541" src="https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-1024x541.png" alt="Load hex firmware from a path in SimulIDE" class="wp-image-39228" srcset="https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-1024x541.png 1024w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-300x159.png 300w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-768x406.png 768w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-1536x812.png 1536w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-2048x1082.png 2048w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-720x380.png 720w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-520x275.png 520w, https://thejeshgn.com/wp-content/uploads/2007/06/simulIDE_loading_firmware_from_a_path-320x169.png 320w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Load hex firmware from a path in SimulIDE</figcaption></figure>
</div>


<p>Now everything is set up. Code and Compile in Arduino IDE. Come to SimulIDE and restart the simulation. That will load the recent firmware and start the simulation.</p>



<h3 class="wp-block-heading">Configure SimulIDE with Arduino Toolchain</h3>



<p>For this, you can use the SimulIDE flatpak version, but you will need the native Arduino installed on the machine. For me, I downloaded the Arduino Linux ZIP file version. Then extracted it into a local directory, say <code>/home/thej/code/arduino-ide_2.3.10_Linux_64bit</code>. Then go to Flatseal and grant SimulIDE access to the folder <code>/home/thej/code/arduino-ide_2.3.10_Linux_64bit</code> .</p>



<p>Once this is done, open an INO file (the Arduino Sketch) on the right-hand panel of the SimulIDE. Go to compiler settings and set the Arduino installation folder path (<code>/home/thej/code/arduino-ide_2.3.10_Linux_64bit</code>) as the Tool path.</p>


<div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide.png" rel="lightbox[39215]"><img loading="lazy" decoding="async" width="1024" height="673" src="https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-1024x673.png" alt="Setup Arduino Toolchain in SimulIDE" class="wp-image-39230" srcset="https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-1024x673.png 1024w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-300x197.png 300w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-768x505.png 768w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-1536x1009.png 1536w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-2048x1346.png 2048w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-720x473.png 720w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-520x342.png 520w, https://thejeshgn.com/wp-content/uploads/2007/06/setup_arduino_toolchainIin_simulide-320x210.png 320w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Setup Arduino Toolchain in SimulIDE</figcaption></figure>
</div>


<p>Now it&#8217;s all set. Use the integrated editor in the SimulIDE to edit code, compile, and load firmware. Then simulate.</p>



<p>Both setups were easy, even though Flatpak made it a bit difficult. If you are not using Flatpak, then you can ignore the Flatseal permission-giving part. The rest of the configuration will still work. Once set up, both flows worked flawlessly. For most people option 2, setting up Arduino toolchain is an easy option. See the demo below.</p>



<figure class="wp-block-video aligncenter"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://thejeshgn.com/wp-content/uploads/2026/07/simulide_arduino_blink_circuit_simulation.webm"></video><figcaption class="wp-element-caption">Demo of Arduino Uno Blink circuit simulation in  SimulIDE </figcaption></figure>



<p>SimulIDE is an amazing piece of software for beginners. You will have fun building and simulating circuits in it. Try it and let me know.</p>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39215" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>


<div class='footnotes' id='footnotes-39215'><div class='footnotedivider'></div><ol><li id='fn-39215-1'> Flatseal is a graphical utility to review and modify permissions from your Flatpak applications. <span class='footnotereverse'><a href='#fnref-39215-1'>&#8617;</a></span></li></ol></div>]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/07/03/getting-started-with-simulide-and-arduino/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		<enclosure url="https://thejeshgn.com/wp-content/uploads/2026/07/simulide_arduino_blink_circuit_simulation.webm" length="8242214" type="video/webm" />

		<post-id xmlns="com-wordpress:feed-additions:1">39215</post-id>	</item>
		<item>
		<title>Building kannada-kasturi-embeddings</title>
		<link>https://thejeshgn.com/2026/07/01/building-kannada-kasturi-embeddings/</link>
					<comments>https://thejeshgn.com/2026/07/01/building-kannada-kasturi-embeddings/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Wed, 01 Jul 2026 11:37:26 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[Digital Kannada]]></category>
		<category><![CDATA[Embedding Models]]></category>
		<category><![CDATA[Free and Open Source]]></category>
		<category><![CDATA[Indic Models]]></category>
		<category><![CDATA[Machine Learning]]></category>
		<category><![CDATA[Sanchaya]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39191</guid>

					<description><![CDATA[Embeddings are numerical representations of real-world objects, such as words, phrases, text, images, audio, and video. Since the real world is so complex, these representations are usually vector arrays of floating-point numbers. This helps computers process meaning, context, and semantic relationships using distances and directions between vectors. Word2Vec and FastText Embeddings are learned&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>Embeddings are numerical representations of real-world objects, such as words, phrases, text, images, audio, and video. Since the real world is so complex, these representations are usually vector arrays of floating-point numbers. This helps computers process meaning, context, and semantic relationships using distances and directions between vectors.</p>






<h3 class="wp-block-heading">Word2Vec and FastText</h3>



<p>Embeddings are learned numerical representations. The actual values depend on the model architecture, training data, and its training objective. So the same object (say a word) gets different embeddings in Word2Vec and GPT. Think of an embedding as a model&#8217;s custom learned coordinate system for meaning.</p>



<p>This coordinate system lets us do useful things. Find semantically similar items, power search, feed downstream ML models, or even perform arithmetic on meaning. A general example from Word2Vec is <code>king − man + woman ≈ queen</code>. There are mainly two types</p>



<ul class="wp-block-list">
<li>SkipGram: Where the model predicts context words given a target word</li>



<li>CBOW (Continuous Bag of Words): Where the model predicts the target word from its context.</li>
</ul>



<p>I wanted to build a simple Word2Vec model for Kannada just to try it out and see if it&#8217;s useful. In the process I also built a <a href="https://fasttext.cc/">FastText</a> model. FastText is the same as Word2Vec with character n-grams. By default, I have built Skipgram based models. But you can try CBOW too.</p>



<p>With n-grams support, FastText makes it easy to handle out of vocabulary words and capture subword information, which Word2vec is not good at. For example, given the settings minn=3, maxn=8, FastText breaks each word into character n-grams of lengths 3 to 8 and learns embeddings for those n-grams as well. There is also a wordNgrams setting, set to 2 by default. We can play with this setting. wordNgrams is used to learn embeddings about the sequences of adjacent words.</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: bash; title: ; notranslate">
For example consider the sentence:

&quot;ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ&quot;

With:

wordNgrams=1 (default) Learns only individual words:

ಬೆಂಗಳೂರು
ಕರ್ನಾಟಕದ
ರಾಜಧಾನಿ

wordNgrams=2 Also learns bigrams:

ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ
ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ

wordNgrams=3 Also learns trigrams:

ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ
</pre></div>


<h3 class="wp-block-heading">Data</h3>



<p>For the corpus, I went with the Kannada text of Kasturi Magazine OCRed by <a href="https://sanchaya.org/" target="_blank" rel="noreferrer noopener">Sanchaya</a> and available on <a href="https://archive.org/search?query=subject%3A%22Kasturi+Magazine%22" target="_blank" rel="noreferrer noopener">Archive.org</a>. My repository has this text; you can use it instead of redownloading from Archive.org. I chose <a href="https://en.wikipedia.org/wiki/Kasthuri_(magazine)" target="_blank" rel="noreferrer noopener">Kasturi</a> content because it&#8217;s the most generic Kannada content <a href="https://archive.org/search?query=subject%3A%22Kasturi+Magazine%22" target="_blank" rel="noreferrer noopener">available</a>, small but good enough to do quick tests. It&#8217;s a monthly family magazine published from the house of Samyukta Karnataka, a very popular newspaper. You can think of it as Reader&#8217;s Digest of Kannada. I have not given the script used to download, as it&#8217;s not required. But it&#8217;s not that difficult to write if you want to. My bots were identified by <code>headers={"User-Agent": "thejeshgn.com/1.0"}</code> and all the downloaded files are named <code>kananda_kasturi_{identifier}.txt</code> where identifier is archive..org identifier.</p>



<h3 class="wp-block-heading">Build and Test</h3>



<p>You can run the trainer, change parameters and test them. I have included all the code and steps for both Fastext and Word2Vec. </p>



<p class="has-text-align-center has-luminous-vivid-amber-background-color has-background">Project is available on <a href="https://codeberg.org/thejeshgn/kannada-kasturi-embeddings">Codeberg</a>, <a href="https://github.com/thejeshgn/kannada-kasturi-embeddings">Github</a> and <a href="https://huggingface.co/thejeshgn/kannada-kasturi-embeddings">HF</a>.</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: bash; title: ; notranslate">
# Prepare the clean corpus.txt from the raw inputs
uv run prepare_corpus.py prepare ../data/ ./corpus.txt

# Train the embeddings with default parameters 
uv run fasttext_embeddings.py train ./corpus.txt --recursive --out ../models/kasturi_fasttext_v4

# Try the embeddings
uv run fasttext_embeddings.py test ../models/kasturi_fasttext_v4.bin --word ಕನ್ನಡ
uv run fasttext_embeddings.py test ../models/kasturi_fasttext_v4.bin --word ರಾಜಧಾನಿ
uv run fasttext_embeddings.py test ../models/kasturi_fasttext_v4.bin --sentence &quot;ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ&quot;
uv run fasttext_embeddings.py compare ../models/kasturi_fasttext_v4.bin
</pre></div>


<p>For comparison and evaluation, I used the same method that I had used before for testing <a href="https://thejeshgn.com/2025/06/18/embedding-models-for-kannada/">other embeddings</a>. Take a bunch of Kannada words (I have used the same bunch), get the cosine similarity between them and create a comparison matrix. I also compared my matrix against the <a href="https://fasttext.cc/docs/en/crawl-vectors.html" target="_blank" rel="noreferrer noopener">word vectors trained by FastText team</a> using Wikipedia and Common Crawl data. You can see them below.</p>



<div class="wp-block-columns is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow"><div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi.png" rel="lightbox[39191]"><img loading="lazy" decoding="async" width="1024" height="914" src="https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-1024x914.png" alt="Compare Matrix for Kannada Kasturi." class="wp-image-39194" srcset="https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-1024x914.png 1024w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-300x268.png 300w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-768x686.png 768w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-1536x1371.png 1536w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-720x643.png 720w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-520x464.png 520w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi-320x286.png 320w, https://thejeshgn.com/wp-content/uploads/2014/03/compare_matrix_kannada_kasturi.png 1593w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Compare Matrix for Kannada Kasturi.</figcaption></figure>
</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow"><div class="wp-block-image">
<figure class="aligncenter size-large"><a href="https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300.png" rel="lightbox[39191]"><img loading="lazy" decoding="async" width="1024" height="914" src="https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-1024x914.png" alt="Kannada Word vectors trained by FastText team - cc.kn.300.bin" class="wp-image-39195" srcset="https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-1024x914.png 1024w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-300x268.png 300w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-768x686.png 768w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-1536x1371.png 1536w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-720x643.png 720w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-520x464.png 520w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300-320x286.png 320w, https://thejeshgn.com/wp-content/uploads/2026/06/compare_matrix_cc_kn_300.png 1593w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Kannada Word vectors trained by FastText team &#8211; cc.kn.300.bin</figcaption></figure>
</div></div>
</div>



<h3 class="wp-block-heading">TODO</h3>



<p>I have called it <em><span style="text-decoration: underline;">kannada-kasturi-embeddings</span></em> because it&#8217;s a good name and we started with Kasturi. But it wont be limited to Kasturi magazine data alone. I will add other newspaper and magazine data to it and build a comprehensive embedding that&#8217;s useful. I will also keep comparing to see how it performs. The current test and compare methods are very basic. I will come up with a better method. So the TODO list is</p>



<ul class="wp-block-list">
<li>Add more quality data</li>



<li>Improve the <code>clean_text</code> function </li>



<li>Add more automated quality testing and comparisons</li>



<li>Build different embedding models and not just Word2Vec or Fastext</li>
</ul>



<p>In the meantime, try it and let me know what you think.</p>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39191" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/07/01/building-kannada-kasturi-embeddings/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">39191</post-id>	</item>
		<item>
		<title>Thank you for attending Back to Basics: Build Your Own LLM from Scratch</title>
		<link>https://thejeshgn.com/2026/06/18/thank-you-for-attending-back-to-basics-build-your-own-llm-from-scratch/</link>
					<comments>https://thejeshgn.com/2026/06/18/thank-you-for-attending-back-to-basics-build-your-own-llm-from-scratch/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Thu, 18 Jun 2026 12:34:19 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[Large Language Model]]></category>
		<category><![CDATA[Study at IITM]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39075</guid>

					<description><![CDATA[I sent this email to all workshop attendees. It made sense to publish it as a post as well. Thank you for attending the &#8220;Back to Basics: Build Your Own LLM from Scratch&#8221; session. It was great to see so much curiosity and many questions in the room, the feedback many of you&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>I sent this email to all workshop attendees. It made sense to publish it as a post as well.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="761" src="https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-1024x761.png" alt="" class="wp-image-39076" srcset="https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-1024x761.png 1024w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-300x223.png 300w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-768x570.png 768w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-1536x1141.png 1536w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-720x535.png 720w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-520x386.png 520w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster-320x238.png 320w, https://thejeshgn.com/wp-content/uploads/2007/07/workshop_poster.png 1800w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p></p>



<p>Thank you for attending the &#8220;<a href="https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/" rel="noreferrer noopener" target="_blank">Back to Basics: Build Your Own LLM from Scratch</a>&#8221; session. It was great to see so much curiosity and many questions in the room, the feedback many of you shared echoed this sentiment. I&#8217;m really glad you found it useful.&nbsp;</p>



<p>As we discussed during the session, I have put together a detailed blog post that walks through everything we covered, including the slides. You can also find the code from the sessions on GitHub and Codeberg.</p>



<ol class="wp-block-list">
<li>Blog with slides and code: <a href="https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/" target="_blank" rel="noreferrer noopener">https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/</a></li>



<li>GitHub: <a href="https://github.com/thejeshgn/workshop-back-basics-llm-build-scratch" target="_blank" rel="noreferrer noopener">https://github.com/thejeshgn/workshop-back-basics-llm-build-scratch</a></li>



<li>Codeberg: <a href="https://codeberg.org/thejeshgn/workshop-back-basics-llm-build-scratch" target="_blank" rel="noreferrer noopener">https://codeberg.org/thejeshgn/workshop-back-basics-llm-build-scratch</a></li>
</ol>



<p>I would also like all of you to continue building on what you learned. You could do that by just forking the code above and adding more features. Some ideas to improve are</p>



<ol class="wp-block-list">
<li>Swap the simple character tokenizer we used for a word-level tokenizer, or go further and implement a BERT-style WordPiece/subword tokenizer, then compare vocabulary size and how each handles unseen words. Also, what effect does it have on the model?</li>



<li>Train on a relatively larger corpus rather than the very small sample we used in the workshop. Project Gutenberg&#8217;s <a href="https://www.gutenberg.org/browse/scores/top" target="_blank" rel="noreferrer noopener">Top 100</a> is a great source of popular, public-domain books that are not very large.</li>



<li>You can also use thematic corpora, such as public-domain poetry collections, recipe collections, or the collected works of Gandhi or Ambedkar, and see how the model&#8217;s generated text picks up the style and vocabulary of that specific body of work. For some of it, you will also have to write data cleaners. Also, be mindful when downloading external content; make sure you don&#8217;t overload their servers and use public domain content. Check whether they already provide downloads in BitTorrent or in other formats, instead of scraping first.</li>



<li>Experiment with model parameters (in GPTConfig) such as context length, heads, and layers once your pipeline works end to end. See how changes in parameters change the model and its output.</li>
</ol>



<p>If you build something interesting, feel free to reach out. I&#8217;d love to hear what you create. Thanks again for being part of the session.</p>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39075" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/06/18/thank-you-for-attending-back-to-basics-build-your-own-llm-from-scratch/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">39075</post-id>	</item>
		<item>
		<title>Back to Basics: Build Your Own LLM from Scratch</title>
		<link>https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/</link>
					<comments>https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Sat, 13 Jun 2026 19:06:52 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[AI 🤖]]></category>
		<category><![CDATA[Assisted by AI 🤖]]></category>
		<category><![CDATA[IITM Paradox]]></category>
		<category><![CDATA[Large Language Model]]></category>
		<category><![CDATA[Study at IITM]]></category>
		<location><![CDATA[Chennai]]></location>
		<guid isPermaLink="false">https://thejeshgn.com/?p=39026</guid>

					<description><![CDATA[I did a workshop titled &#8220;Back to Basics: Build Your Own LLM from Scratch&#8221; at IITM/Paradox 2026, which kind of included some basic theory on how a transformer works, and then building a very small LLM. The idea was to demystify an LLM (or transformer) by understanding what goes on and then building&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>I did a workshop titled &#8220;Back to Basics: Build Your Own LLM from Scratch&#8221; at <a href="https://thejeshgn.com/tag/iitm-paradox/?order=asc">IITM/Paradox 2026</a>, which kind of included some basic theory on how a transformer works, and then building a very small LLM. The idea was to demystify an LLM (or transformer) by understanding what goes on and then building one to deepen our understanding. I had to skip some slides because the planned session was only two hours. Ideally, I want it to be around 4 hours, split into 2 sessions: one for theory and one for lab. Maybe next time, when I plan, I will make it 4 hours so I can do it at a slower pace.</p>



<p>Of course, there are other similar workshops available online, and some of them are linked in the references section. This is just my take on it and what I used for my own understanding.</p>



<p>Suppose you want to try it at your own pace. Try the <a href="https://thejeshgn.com/wp-content/uploads/2026/06/back-basics-llm-build-scratch.html">slides</a> below and then use annotated code to read and run. Slides and code are in a repo (<a href="https://codeberg.org/thejeshgn/workshop-back-basics-llm-build-scratch">CB</a>, <a href="https://github.com/thejeshgn/workshop-back-basics-llm-build-scratch">GH</a>) too if you prefer that.</p>



<iframe loading="lazy" width="99%" height="550px" src="https://thejeshgn.com/wp-content/uploads/2026/06/back-basics-llm-build-scratch.html" scrolling="no" frameBorder="0"></iframe>



<p></p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="766" height="927" src="https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram.png" alt="" class="wp-image-39038" srcset="https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram.png 766w, https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram-248x300.png 248w, https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram-720x871.png 720w, https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram-520x629.png 520w, https://thejeshgn.com/wp-content/uploads/2017/02/character_gpt_workshop_diagram-320x387.png 320w" sizes="auto, (max-width: 766px) 100vw, 766px" /></figure>



<p></p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: python; title: ; notranslate">
#!/usr/bin/env -S uv run --script
# /// script
# requires-python = &quot;&gt;=3.10&quot;
# dependencies = &#x5B;
#   &quot;torch==2.8.0&quot;,
# ]
# &#x5B;tool.uv]
# extra-index-url = &#x5B;&quot;https://download.pytorch.org/whl/cpu&quot;]
# ///
&quot;&quot;&quot;
build_and_test.py  (ANNOTATED VERSION FOR WORKSHOP)
A minimal, single-file GPT for &quot;Back to Basics: Build Your Own LLM from Scratch&quot;.

═══════════════════════════════════════════════════════════════════════════════
WALKTHROUGH MAP — suggested order
═══════════════════════════════════════════════════════════════════════════════

  ① GPTConfig            ← slide &quot;The whole model in 5 numbers&quot;
  ② CharTokenizer        ← slides &quot;Tokens&quot; / &quot;Tokenizer&quot;
  ③ GPT.__init__/forward ← slide &quot;Recap: what we just built&quot; (the big pipeline)
  ④ CausalSelfAttention  ← slides &quot;Single-head attention&quot; → &quot;Combining heads with Wo&quot;
  ⑤ FeedForward          ← slide &quot;Feed-Forward Network (FFN)&quot;
  ⑥ TransformerBlock     ← slide &quot;One full transformer block&quot;
  ⑦ get_batch            ← (where x/y &quot;next-token&quot; pairs come from)
  ⑧ train()              ← slides &quot;Cross-entropy&quot; → &quot;Training&quot;
  ⑨ GPT.generate()       ← slide &quot;Generation: from logits to text&quot;

Comments marked 💬 are things worth SAYING out loud.
Comments marked ❓ are good questions to ASK the room.
Comments marked ⚠️ are common gotchas / likely audience questions.

Usage:
    # Train on a text file (CPU only, by design)
    uv run build_and_test.py train --data ../data/shakespeare.txt --max-steps 2000

    # Generate from a saved checkpoint
    uv run build_and_test.py generate --checkpoint ../checkpoints/run1/final_checkpoint.pt --prompt &quot;To be, or not &quot; --num-new-tokens 200 --temperature 0.8 --top-k 40 --seed 42
&quot;&quot;&quot;

import argparse
import csv
import math
import os
import sys
import time
from dataclasses import dataclass, asdict

import torch
import torch.nn as nn
import torch.nn.functional as F


# ═════════════════════════════════════════════════════════════════════════════
# ① CONFIG                                  &#x5B;Slide: &quot;The whole model in 5 numbers&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;These five numbers ARE the model. Everything else is derived from them.&quot;
#    Point back to this class every time a shape like (B, T, C) appears below:
#       B = batch_size, T ≤ block_size, C = n_embd.

@dataclass
class GPTConfig:
    # Architecture (the 5 numbers from the slides)
    vocab_size: int = 65        # how many unique tokens (set from data, after tokenizer)
    n_embd:     int = 256       # C: the vector size each token is represented by
    n_head:     int = 4         # parallel attention heads (d_k = n_embd/n_head = 64)
    block_size: int = 256       # T_max: the longest context the model can ever see
    n_layer:    int = 4         # how many TransformerBlocks we stack

# Training knobs — deliberately NOT in GPTConfig:
# 💬 &quot;These shape the *training run*, not the *model*. A checkpoint doesn&#039;t need them.&quot;
batch_size = 32   # sequences per training step (the B in (B, T, C))
dropout = 0.1     # ⚠️ on during training, automatically off in eval()
                  #    &#x5B;Slide: &quot;Training vs. inference&quot; — dropout row]

output_path = &quot;../checkpoints/run1&quot;

# ═════════════════════════════════════════════════════════════════════════════
# ② TOKENIZER                                    &#x5B;Slides: &quot;Tokens&quot; / &quot;Tokenizer&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;A tokenizer has exactly two jobs: encode (text → IDs) and decode (IDs → text).
#    We picked character-level — the simplest of the three choices on the slide.
#    GPT/LLaMA/Claude use subword BPE; same idea, fancier vocab.&quot;
#
# ❓ Ask: &quot;If our vocab is the 65 unique characters in Shakespeare, what happens
#    when you prompt with an emoji?&quot; → see encode(): it silently drops unknowns.

class CharTokenizer:
    &quot;&quot;&quot;Smallest possible tokenizer: one character = one token.&quot;&quot;&quot;

    def __init__(self, vocab: list&#x5B;str]):
        self.vocab = vocab
        # stoi = &quot;string to int&quot;, itos = &quot;int to string&quot; — two dicts, that&#039;s it.
        self.stoi = {ch: i for i, ch in enumerate(vocab)}
        self.itos = {i: ch for i, ch in enumerate(vocab)}

    @classmethod
    def from_text(cls, text: str) -&gt; &quot;CharTokenizer&quot;:
        # 💬 &quot;The vocab is just every unique character that occurs in the data.&quot;
        # Sorted so the vocab is deterministic across runs
        # ⚠️ Without sorted(), set() ordering varies → token IDs change between runs
        #    → an old checkpoint would decode to garbage. This one line is why we
        #    can reload checkpoints reliably.
        vocab = sorted(list(set(text)))
        return cls(vocab)

    def encode(self, s: str) -&gt; list&#x5B;int]:
        # text → list of integers.  &#x5B;Slide: &quot;Tokenizer&quot; — encode example]
        # `if c in self.stoi`: characters not in the training data are dropped.
        return &#x5B;self.stoi&#x5B;c] for c in s if c in self.stoi]

    def decode(self, ids: list&#x5B;int]) -&gt; str:
        # integers → text. Perfect inverse of encode (for known chars).
        return &quot;&quot;.join(self.itos&#x5B;i] for i in ids)

    @property
    def vocab_size(self) -&gt; int:
        # 💬 &quot;This becomes the first of our 5 numbers — vocab_size in GPTConfig.&quot;
        return len(self.vocab)


# ═════════════════════════════════════════════════════════════════════════════
# ④ ATTENTION             &#x5B;Slides: &quot;Why attention?&quot; → &quot;Combining heads with Wo&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;This class is the heart of the workshop. The 7 numbered steps in forward()
#    map one-to-one onto the attention slides. Everything else is plumbing.&quot;
#
# Teaching tip: walk forward() with a concrete shape, e.g. B=32, T=256, C=256,
# n_head=4, d_k=64 — and write the shapes on the board as you go.

class CausalSelfAttention(nn.Module):
    &quot;&quot;&quot;Multi-head causal self-attention. One linear projects to Q,K,V together.&quot;&quot;&quot;

    def __init__(self, cfg: GPTConfig):
        super().__init__()
        # d_k = n_embd / n_head must divide evenly — each head gets a clean slice.
        # &#x5B;Slide: &quot;Multi-head attention&quot; — d_k = n_embd / n_head]
        assert cfg.n_embd % cfg.n_head == 0, &quot;n_embd must be divisible by n_head&quot;
        self.n_head = cfg.n_head
        self.n_embd = cfg.n_embd
        self.d_k = cfg.n_embd // cfg.n_head

        # 💬 &quot;The slides show three separate matrices Wq, Wk, Wv, each (C × C).
        #    In code we fuse them into ONE (C × 3C) matrix for efficiency —
        #    one matmul instead of three. Same math, same parameter count.&quot;
        # &#x5B;Slide: &quot;Single-head attention: Q, K, V&quot;]
        self.qkv = nn.Linear(cfg.n_embd, 3 * cfg.n_embd)

        # Wo from the slides: lets the heads talk to each other after running
        # independently. Without it, heads would be siloed.
        # &#x5B;Slide: &quot;Combining heads with Wo&quot;]
        self.proj = nn.Linear(cfg.n_embd, cfg.n_embd)
        self.dropout = nn.Dropout(dropout)

        # 💬 &quot;The causal mask is NOT learned — it&#039;s a fixed triangle of 1s.
        #    Row i has 1s up to column i: &#039;token i may look at tokens 0..i&#039;.&quot;
        # &#x5B;Slide: &quot;Causal mask&quot;] and &#x5B;Slide: &quot;What gets learned, what stays fixed&quot;]
        # register_buffer = &quot;part of the model, moves with .to(device),
        # saved in state_dict, but NO gradients&quot; — perfect for a constant.
        mask = torch.tril(torch.ones(cfg.block_size, cfg.block_size))
        self.register_buffer(&quot;mask&quot;, mask.view(1, 1, cfg.block_size, cfg.block_size))

    def forward(self, x):
        B, T, C = x.shape  # batch, seq_len, n_embd — write these on the board

        # ── 1) Project to Q, K, V ──────────────── &#x5B;Slide: &quot;Q, K, V&quot;]
        # One big matmul gives (B, T, 3C); split() carves it into three (B, T, C).
        # 💬 &quot;Query: what am I looking for? Key: what do I offer? Value: what do
        #    I pass along if matched?&quot;
        q, k, v = self.qkv(x).split(self.n_embd, dim=2)

        # ── 2) Split into heads ────────── &#x5B;Slide: &quot;Multi-head attention&quot;]
        # (B, T, C) → (B, T, n_head, d_k) → transpose → (B, n_head, T, d_k)
        # 💬 &quot;No new computation here — we&#039;re just reshaping so each head can run
        #    the SAME attention math independently on its own d_k-sized slice.&quot;
        q = q.view(B, T, self.n_head, self.d_k).transpose(1, 2)
        k = k.view(B, T, self.n_head, self.d_k).transpose(1, 2)
        v = v.view(B, T, self.n_head, self.d_k).transpose(1, 2)

        # ── 3) Scaled dot-product scores ──── &#x5B;Slide: &quot;Attention scores (scaled)&quot;]
        # (B, nh, T, d_k) @ (B, nh, d_k, T) → (B, nh, T, T)
        # 💬 &quot;A T×T grid per head: how relevant is every token to every other token.&quot;
        # ⚠️ The 1/√d_k scaling is the line students forget. Without it, dot
        #    products grow with d_k → softmax saturates → gradients vanish.
        scores = (q @ k.transpose(-2, -1)) * (1.0 / math.sqrt(self.d_k))

        # ── 4) Causal mask ──────────────────────── &#x5B;Slide: &quot;Causal mask&quot;]
        # Where the triangle has 0 (future positions), drop in -inf.
        # 💬 &quot;-inf BEFORE softmax becomes exactly 0 AFTER softmax — the model
        #    literally cannot peek at the answer.&quot;
        # &#x5B;:T, :T] crops the precomputed block_size mask to the actual seq length.
        scores = scores.masked_fill(self.mask&#x5B;:, :, :T, :T] == 0, float(&quot;-inf&quot;))

        # ── 5) Softmax → attention weights ──── &#x5B;Slide: &quot;Softmax intuition&quot;]
        # Each row becomes a probability distribution: positive, sums to 1,
        # behaves like &quot;importance&quot;.
        attn = F.softmax(scores, dim=-1)
        attn = self.dropout(attn)  # regularization: randomly drop some attention links

        # ── 6) Apply attention to V ────────────── &#x5B;Slide: &quot;Apply attention&quot;]
        # (B, nh, T, T) @ (B, nh, T, d_k) → (B, nh, T, d_k)
        # 💬 &quot;Each output row is a weighted BLEND of value vectors from earlier
        #    positions. THIS is the heart of the transformer.&quot;
        out = attn @ v

        # ── 7) Re-combine heads + Wo ──── &#x5B;Slide: &quot;Combining heads with Wo&quot;]
        # (B, nh, T, d_k) → (B, T, nh, d_k) → (B, T, C): concat of all heads.
        # ⚠️ .contiguous() is needed because transpose only changes the view,
        #    not memory layout — .view() requires contiguous memory.
        out = out.transpose(1, 2).contiguous().view(B, T, C)
        out = self.proj(out)  # Wo: mixes information across heads
        return out


# ═════════════════════════════════════════════════════════════════════════════
# ⑤ FFN                                  &#x5B;Slide: &quot;Feed-Forward Network (FFN)&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;Attention mixes information ACROSS tokens; the FFN processes each token
#    INDEPENDENTLY — same two layers applied to every position. Expand to 4×
#    (room to think), non-linearity, compress back.&quot;

class FeedForward(nn.Module):
    &quot;&quot;&quot;Two-layer MLP: expand to 4x, GELU, compress back.&quot;&quot;&quot;

    def __init__(self, cfg: GPTConfig):
        super().__init__()
        d_ff = 4 * cfg.n_embd            # the classic 4× expansion (slide: d_ff = 4d)
        self.fc1 = nn.Linear(cfg.n_embd, d_ff)   # W1: expand  (C → 4C)
        self.fc2 = nn.Linear(d_ff, cfg.n_embd)   # W2: compress (4C → C)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        # expand → GELU (GPT&#039;s choice over ReLU) → compress → dropout
        # ❓ Ask: &quot;Why is the non-linearity essential?&quot; → without it, fc2(fc1(x))
        #    collapses into a single linear layer; depth would buy nothing.
        return self.dropout(self.fc2(F.gelu(self.fc1(x))))


# ═════════════════════════════════════════════════════════════════════════════
# ⑥ TRANSFORMER BLOCK                  &#x5B;Slide: &quot;One full transformer block&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;This is the repeating unit — the slide diagram in 3 lines of code.
#    PRE-norm: LayerNorm goes BEFORE each sublayer (GPT-2/modern convention),
#    and the residual &#039;x +&#039; is the highway that lets gradients flow through
#    deep stacks.&quot;
# &#x5B;Slide: &quot;Residual + LayerNorm&quot;]

class TransformerBlock(nn.Module):
    &quot;&quot;&quot;Pre-norm block: LN -&gt; Attn -&gt; +residual -&gt; LN -&gt; FFN -&gt; +residual&quot;&quot;&quot;

    def __init__(self, cfg: GPTConfig):
        super().__init__()
        self.ln1 = nn.LayerNorm(cfg.n_embd)   # γ, β — learned (2 × n_embd params)
        self.attn = CausalSelfAttention(cfg)
        self.ln2 = nn.LayerNorm(cfg.n_embd)   # second LN, own γ, β
        self.ffn = FeedForward(cfg)

    def forward(self, x):
        # 💬 Read these aloud as: &quot;x plus attention-of-normalized-x&quot; —
        #    the residual means each sublayer only learns a CORRECTION to x.
        x = x + self.attn(self.ln1(x))   # sublayer 1: communicate (across tokens)
        x = x + self.ffn(self.ln2(x))    # sublayer 2: compute (per token)
        return x


# ═════════════════════════════════════════════════════════════════════════════
# ③ THE FULL GPT                          &#x5B;Slide: &quot;Recap: what we just built&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 Teaching tip: show __init__ + forward() FIRST as the bird&#039;s-eye view —
#    it mirrors the recap-slide pipeline line by line — then descend into
#    attention. People hold details better once they&#039;ve seen the skeleton.

class GPT(nn.Module):
    &quot;&quot;&quot;The whole model: embeddings + N blocks + final LN + LM head.&quot;&quot;&quot;

    def __init__(self, cfg: GPTConfig):
        super().__init__()
        self.cfg = cfg

        # Token embedding table: (vocab_size × n_embd), learned lookup.
        # &#x5B;Slide: &quot;Embeddings&quot; — &quot;initialized randomly, updated during training&quot;]
        self.tok_emb = nn.Embedding(cfg.vocab_size, cfg.n_embd)

        # LEARNED positional encoding (GPT-style), one vector per position.
        # &#x5B;Slide: &quot;Positional encoding&quot; — the &#039;Learned&#039; flavor, not sinusoidal]
        self.pos_emb = nn.Embedding(cfg.block_size, cfg.n_embd)

        self.drop = nn.Dropout(dropout)

        # The stack: n_layer identical-shaped blocks, each with its OWN weights.
        # &#x5B;Slide: &quot;Stacking layers&quot;]
        self.blocks = nn.ModuleList(&#x5B;TransformerBlock(cfg) for _ in range(cfg.n_layer)])

        self.ln_f = nn.LayerNorm(cfg.n_embd)  # final LN before the head

        # LM head: project (B, T, C) back to (B, T, vocab) — a score per token.
        # &#x5B;Slide: &quot;Output logits&quot;]
        self.head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)

        # 💬 WEIGHT TYING: the output head and the input embedding SHARE one
        #    matrix (used transposed in the matmul). Saves vocab_size × n_embd
        #    params and works well in practice — mentioned on the logits slide.
        # ⚠️ This is why the parameter count printout is ~16K lower than the
        #    worked example on the slides (which counts the head separately).
        self.head.weight = self.tok_emb.weight

        # Initialize all weights small and Gaussian (std 0.02, the GPT-2 recipe).
        # ❓ &quot;Why not zeros?&quot; → all-zero weights = all neurons identical = no
        #    symmetry breaking; nothing distinct to learn.
        self.apply(self._init_weights)

    def _init_weights(self, m):
        if isinstance(m, nn.Linear):
            nn.init.normal_(m.weight, mean=0.0, std=0.02)
            if m.bias is not None:
                nn.init.zeros_(m.bias)
        elif isinstance(m, nn.Embedding):
            nn.init.normal_(m.weight, mean=0.0, std=0.02)

    def num_parameters(self) -&gt; int:
        # 💬 Compare the printout with the slide&#039;s worked example (~3.25M).
        # &#x5B;Slide: &quot;Parameter count: worked example&quot;]
        return sum(p.numel() for p in self.parameters() if p.requires_grad)

    def forward(self, idx, targets=None):
        &quot;&quot;&quot;
        idx: (B, T) token IDs
        targets: (B, T) next-token IDs (for training); None for inference
        returns: logits (B, T, vocab_size), loss or None

        💬 &quot;ONE forward pass serves both training and inference — same brain,
           different loop. targets=None is the only switch.&quot;
        &#x5B;Slide: &quot;Training vs. inference&quot;]
        &quot;&quot;&quot;
        B, T = idx.shape
        assert T &lt;= self.cfg.block_size, f&quot;sequence length {T} &gt; block_size {self.cfg.block_size}&quot;

        # ── The recap-slide pipeline, line by line ──
        tok = self.tok_emb(idx)                                       # (B, T, C)  token IDs → vectors
        pos = self.pos_emb(torch.arange(T, device=idx.device))        # (T, C)     position 0..T-1 → vectors
        x = self.drop(tok + pos)                                      # (B, T, C)  X = E + position
        # ⚠️ tok is (B,T,C), pos is (T,C) — broadcasting adds the same position
        #    vectors to every sequence in the batch. Worth pausing on.

        for block in self.blocks:        # n_layer blocks, each refines x
            x = block(x)
        x = self.ln_f(x)                 # final LayerNorm  &#x5B;Slide: &quot;Stacking layers&quot;]
        logits = self.head(x)                                         # (B, T, vocab)

        loss = None
        if targets is not None:
            # ── TRAINING branch ──        &#x5B;Slides: &quot;Cross-entropy loss&quot;]
            # 💬 &quot;The model predicts the next token at EVERY position in
            #    parallel — T predictions per sequence, not 1. That&#039;s why
            #    transformer training is so efficient.&quot;
            # Flatten (B, T, vocab) → (B·T, vocab) and (B, T) → (B·T,)
            # because F.cross_entropy wants (N, classes) and (N,) of true IDs.
            # ⚠️ cross_entropy takes raw LOGITS — it applies softmax + -log(p_t)
            #    internally. Don&#039;t softmax twice (a classic live-coding bug).
            loss = F.cross_entropy(
                logits.view(-1, logits.size(-1)),
                targets.view(-1),
            )
        return logits, loss

    # ═════════════════════════════════════════════════════════════════════════
    # ⑨ GENERATION                  &#x5B;Slide: &quot;Generation: from logits to text&quot;]
    # ═════════════════════════════════════════════════════════════════════════
    @torch.no_grad()   # inference: no gradients, no backprop — weights frozen
    def generate(self, idx, num_new_tokens: int, temperature: float = 1.0,
                 top_k: int | None = None) -&gt; torch.Tensor:
        &quot;&quot;&quot;Autoregressively generate num_new_tokens tokens.

        💬 &quot;The slide&#039;s loop: last logits → softmax → sample → append → repeat.
           One token at a time. This is how GPT writes a sentence.&quot;
        &quot;&quot;&quot;
        self.eval()  # switches dropout OFF &#x5B;Slide: &quot;Training vs. inference&quot;]
        for _ in range(num_new_tokens):
            # If the running text exceeds block_size, keep only the last
            # block_size tokens — the model can&#039;t attend beyond its context.
            # 💬 &quot;This IS the &#039;context window&#039; people talk about in big LLMs.&quot;
            idx_cond = idx if idx.size(1) &lt;= self.cfg.block_size else idx&#x5B;:, -self.cfg.block_size:]

            logits, _ = self(idx_cond)   # full forward pass; loss is None here

            # Take only the LAST position&#039;s logits — the next-token prediction.
            # (Training used all T positions; inference uses just one.)
            # TEMPERATURE: divide logits before softmax.
            #   &lt;1.0 sharpens (more confident/repetitive), &gt;1.0 flattens (wilder).
            # ⚠️ max(temperature, 1e-8) guards against divide-by-zero at temp=0.
            logits = logits&#x5B;:, -1, :] / max(temperature, 1e-8)

            # TOP-K: keep only the k highest-scoring tokens, set the rest to
            # -inf (so softmax gives them probability 0). Stops the model from
            # ever sampling a wildly unlikely character.
            if top_k is not None and top_k &gt; 0:
                v, _ = torch.topk(logits, min(top_k, logits.size(-1)))
                logits&#x5B;logits &lt; v&#x5B;:, &#x5B;-1]]] = float(&quot;-inf&quot;)  # v&#x5B;:, &#x5B;-1]] = k-th best score

            probs = F.softmax(logits, dim=-1)                 # scores → probabilities
            next_id = torch.multinomial(probs, num_samples=1) # SAMPLE (not argmax) (B, 1)
            # ❓ Ask: &quot;What changes if we use argmax instead?&quot; → deterministic,
            #    and typically loops/repeats. Sampling is where variety comes from.
            idx = torch.cat(&#x5B;idx, next_id], dim=1)            # append &amp; loop
        return idx


# ═════════════════════════════════════════════════════════════════════════════
# ⑦ DATA: making (input, target) pairs
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;Where do the &#039;answer keys&#039; come from? The text itself. The target is the
#    input shifted one character to the right — free labels, no human needed.
#    This is what &#039;self-supervised&#039; means.&quot;
#
#    data:   &#x5B;T, h, e, _, c, a, t]
#    x  =     &#x5B;T, h, e, _, c, a]
#    y  =        &#x5B;h, e, _, c, a, t]   ← y&#x5B;i] is the &#039;next token&#039; after x&#x5B;i]

def get_batch(data: torch.Tensor, block_size: int, batch_size: int,
              device: torch.device) -&gt; tuple&#x5B;torch.Tensor, torch.Tensor]:
    &quot;&quot;&quot;Sample batch_size random windows of length block_size from data.&quot;&quot;&quot;
    # Random start indices; -1 leaves room for the shifted target.
    ix = torch.randint(0, len(data) - block_size - 1, (batch_size,))
    x = torch.stack(&#x5B;data&#x5B;i:i + block_size] for i in ix])          # inputs
    y = torch.stack(&#x5B;data&#x5B;i + 1:i + 1 + block_size] for i in ix])  # same, shifted +1
    return x.to(device), y.to(device)
    # ⚠️ Random windows ≠ epochs. We sample with replacement, so &quot;one epoch&quot;
    #    isn&#039;t well-defined here — we just count steps. Fine at this scale.


# ═════════════════════════════════════════════════════════════════════════════
# ⑧ TRAIN COMMAND        &#x5B;Slides: &quot;Why we minimize it&quot; → &quot;Training&quot;]
# ═════════════════════════════════════════════════════════════════════════════
# 💬 The training slide&#039;s loop, in code:
#    batch → forward → loss → backward (gradients) → optimizer step → repeat.
#    Map each line of the loop below onto that diagram as you scroll.

def train(args):
    device = torch.device(&quot;cpu&quot;)  # CPU-only, by design for the workshop
    torch.manual_seed(1337)       # fixed seed → everyone in the room gets the
                                  # same loss curve. All of us get same model.

    # ── 1) Load data: just a plain text file ──
    if not os.path.exists(args.data):
        sys.exit(f&quot;Data file not found: {args.data}&quot;)
    with open(args.data, &quot;r&quot;, encoding=&quot;utf-8&quot;) as f:
        text = f.read()
    print(f&quot;Loaded {len(text):,} characters from {args.data}&quot;)

    # ── 2) Build tokenizer FROM the data ──
    # 💬 &quot;vocab_size isn&#039;t chosen by us — it falls out of the data. For tiny
    #    Shakespeare it&#039;s 65: letters, digits, punctuation, newline.&quot;
    tokenizer = CharTokenizer.from_text(text)
    print(f&quot;Vocab size: {tokenizer.vocab_size}&quot;)

    # ── 3) Encode the whole corpus ONCE; split 90/10 train/val ──
    # ❓ Ask: &quot;Why hold out a validation set?&quot; → train loss can fall from
    #    memorization; val loss tells us if the model GENERALIZES. Watch the
    #    gap between the two columns in the printout.
    data = torch.tensor(tokenizer.encode(text), dtype=torch.long)
    n_train = int(0.9 * len(data))
    train_data = data&#x5B;:n_train]
    val_data = data&#x5B;n_train:]
    print(f&quot;Train tokens: {len(train_data):,}   Val tokens: {len(val_data):,}&quot;)

    # ── 4) Build the model from the 5 numbers ──
    cfg = GPTConfig(vocab_size=tokenizer.vocab_size)
    model = GPT(cfg).to(device)
    # 💬 Pause on this printout and reconcile it with the parameter-count
    #    slides (~3.25M). Slight difference = weight tying (head not double-counted).
    print(f&quot;Model parameters: {model.num_parameters():,}&quot;)

    # ── 5) Optimizer: AdamW ──   &#x5B;Slide: &quot;From gradients to weight updates&quot;]
    # 💬 &quot;AdamW = the slide&#039;s `w -= lr × gradient`, but with per-weight adaptive
    #    step sizes from running averages of past gradients. Used by GPT-2/3,
    #    LLaMA — and by us.&quot;
    # weight_decay gently pulls weights toward 0 — regularization.
    optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.1)

    # LR schedule: linear WARMUP (first 100 steps), then COSINE DECAY to min_lr.
    # 💬 &quot;Warmup: start gentle while weights are random garbage. Cosine: take
    #    smaller steps as we converge — like slowing down when parallel parking.&quot;
    # ❓ &quot;What&#039;s lr at step 0? At step `warmup`? At step max_steps?&quot; → trace it.
    def lr_at(step: int, max_steps: int, base_lr: float = 3e-4,
              warmup: int = 100, min_lr: float = 3e-5) -&gt; float:
        if step &lt; warmup:
            return base_lr * (step + 1) / warmup            # ramp 0 → base_lr
        progress = (step - warmup) / max(1, max_steps - warmup)
        progress = min(1.0, progress)
        return min_lr + 0.5 * (base_lr - min_lr) * (1 + math.cos(math.pi * progress))

    # ── 6) Training loop ──
    # Fixed sample prompts so the audience can WATCH the same prompt improve
    # from noise → words → Shakespeare-ish as training progresses.
    sample_prompts = &#x5B;&quot;To be, or not &quot;, &quot;For I am falser than vows made in&quot;]
    log_path = f&quot;{output_path}/loss_log.csv&quot;      # for graphing the loss curve later
    log_file = open(log_path, &quot;w&quot;, newline=&quot;&quot;)
    log_writer = csv.writer(log_file)
    log_writer.writerow(&#x5B;&quot;step&quot;, &quot;train_loss&quot;, &quot;val_loss&quot;, &quot;lr&quot;])

    # Derived intervals: 10 evals, 5 sample dumps, 4 checkpoints per run,
    # regardless of --max-steps.
    eval_every = max(1, args.max_steps // 10)
    sample_every = max(1, args.max_steps // 5)
    ckpt_every = max(1, args.max_steps // 4)

    t0 = time.time()
    model.train()  # dropout ON
    for step in range(args.max_steps):
        # Set this step&#039;s learning rate (PyTorch optimizers read it from
        # param_groups; we overwrite it each step with our schedule).
        lr = lr_at(step, args.max_steps)
        for g in optimizer.param_groups:
            g&#x5B;&quot;lr&quot;] = lr

        # ══ THE four lines that ARE training ══  &#x5B;Slide: &quot;Training&quot;]
        xb, yb = get_batch(train_data, cfg.block_size, batch_size, device)
        _, loss = model(xb, yb)               # 1. forward → loss (one number)

        optimizer.zero_grad(set_to_none=True) # 2. clear last step&#039;s gradients
        # ⚠️ Forgetting zero_grad is THE classic bug: PyTorch ACCUMULATES
        #    gradients by default, so they&#039;d pile up across steps.
        loss.backward()                       # 3. backprop: chain rule, automatic
                                              #    &#x5B;Slide: &quot;How backprop computes gradients&quot;]
        optimizer.step()                      # 4. nudge EVERY weight a tiny bit

        # ── Periodic eval + logging ──
        if step % eval_every == 0 or step == args.max_steps - 1:
            model.eval()                      # dropout OFF for a fair measurement
            with torch.no_grad():             # no gradient bookkeeping needed
                xv, yv = get_batch(val_data, cfg.block_size, batch_size, device)
                _, val_loss = model(xv, yv)
            elapsed = time.time() - t0
            # 💬 Narrate the first line: loss ≈ 4.17 ≈ ln(65) — the loss of a
            #    UNIFORM guess over 65 chars (&quot;Model B&quot; on the cross-entropy
            #    slide). Watching it fall below that = the model is learning.
            print(f&quot;step {step:5d} | lr {lr:.2e} | train {loss.item():.4f} | &quot;
                  f&quot;val {val_loss.item():.4f} | {elapsed:.1f}s&quot;)
            log_writer.writerow(&#x5B;step, loss.item(), val_loss.item(), lr])
            log_file.flush()                  # so the CSV is graphable mid-run
            model.train()                     # back to training mode

        # ── Periodic samples: the workshop&#039;s &quot;wow&quot; moment ──
        # 💬 Early samples are gibberish; mid-run grows words and line breaks;
        #    late samples look like a drunk Shakespeare. Same prompts each time
        #    makes the progress visible.
        if step % sample_every == 0 and step &gt; 0:
            model.eval()
            for p in sample_prompts:
                ids = torch.tensor(&#x5B;tokenizer.encode(p)], dtype=torch.long, device=device)
                out = model.generate(ids, num_new_tokens=80, temperature=0.8, top_k=20)
                generated = tokenizer.decode(out&#x5B;0].tolist())
                print(f&quot;   sample: {generated!r}&quot;)
            model.train()

        # ── Periodic checkpoints (resume / compare across training stages) ──
        if step &gt; 0 and step % ckpt_every == 0:
            save_checkpoint(model, tokenizer, cfg, f&quot;{output_path}/ckpt_step_{step}.pt&quot;)

    # Final checkpoint — this is what the generate command loads.
    save_checkpoint(model, tokenizer, cfg, f&quot;{output_path}/final_checkpoint.pt&quot;)
    log_file.close()
    print(f&quot;Done. Wrote final_checkpoint.pt and {log_path}.&quot;)


def save_checkpoint(model: GPT, tokenizer: CharTokenizer, cfg: GPTConfig, path: str):
    # 💬 &quot;A checkpoint must contain everything needed to rebuild the model:
    #    1. the weights, 2. the 5 numbers (shape), 3. the vocab (so token IDs
    #    decode to the same characters). Forget the vocab → garbage output.&quot;
    torch.save({
        &quot;model_state&quot;: model.state_dict(),   # every learned tensor, by name
        &quot;config&quot;: asdict(cfg),               # the 5 numbers
        &quot;vocab&quot;: tokenizer.vocab,            # the character list
    }, path)
    print(f&quot;  saved {path}&quot;)


# ═════════════════════════════════════════════════════════════════════════════
# GENERATE COMMAND — checkpoint in, text out
# ═════════════════════════════════════════════════════════════════════════════
# 💬 &quot;Inference = rebuild the exact same model, load frozen weights, loop
#    generate(). No loss, no gradients, no optimizer — compare the two columns
#    of the &#039;Training vs. inference&#039; slide.&quot;

def generate(args):
    device = torch.device(&quot;cpu&quot;)
    if args.seed is not None:
        torch.manual_seed(args.seed)  # same seed + same prompt → same output
        # 💬 Good demo: run twice with --seed 42 (identical), then without (varies).

    if not os.path.exists(args.checkpoint):
        sys.exit(f&quot;Checkpoint not found: {args.checkpoint}&quot;)
    # ⚠️ weights_only=False because our checkpoint also carries config + vocab
    #    (not just tensors). Fine for our OWN files; for untrusted downloads
    #    you&#039;d want weights_only=True (it restricts unpickling).
    ckpt = torch.load(args.checkpoint, map_location=device, weights_only=False)

    # Rebuild the exact architecture and tokenizer the checkpoint was saved with:
    cfg = GPTConfig(**ckpt&#x5B;&quot;config&quot;])        # the 5 numbers → same shapes
    tokenizer = CharTokenizer(ckpt&#x5B;&quot;vocab&quot;]) # same vocab → same ID↔char mapping

    model = GPT(cfg).to(device)
    model.load_state_dict(ckpt&#x5B;&quot;model_state&quot;])  # pour the learned weights back in
    model.eval()                                # inference mode: dropout off

    # prompt → IDs → generate → IDs → text. The full round trip from slide 1.
    ids = torch.tensor(&#x5B;tokenizer.encode(args.prompt)], dtype=torch.long, device=device)
    out = model.generate(
        ids,
        num_new_tokens=args.num_new_tokens,
        temperature=args.temperature,   # ❓ live demo: try 0.2 vs 1.5 and compare
        top_k=args.top_k,
    )
    print(tokenizer.decode(out&#x5B;0].tolist()))


# ═════════════════════════════════════════════════════════════════════════════
# CLI — two subcommands, as promised on the &quot;Hands-on: the plan&quot; slide
# ═════════════════════════════════════════════════════════════════════════════

def main():
    parser = argparse.ArgumentParser(description=&quot;Tiny GPT: train and generate&quot;)
    sub = parser.add_subparsers(dest=&quot;cmd&quot;, required=True)

    # train: only data + steps are CLI args; architecture lives in GPTConfig.
    p_train = sub.add_parser(&quot;train&quot;, help=&quot;Train the model from a text file&quot;)
    p_train.add_argument(&quot;--data&quot;, required=True, help=&quot;path to UTF-8 text file&quot;)
    p_train.add_argument(&quot;--max-steps&quot;, type=int, default=2000)
    p_train.set_defaults(func=train)

    # generate: checkpoint + prompt + the sampling knobs from the slides.
    p_gen = sub.add_parser(&quot;generate&quot;, help=&quot;Generate text from a checkpoint&quot;)
    p_gen.add_argument(&quot;--checkpoint&quot;, required=True)
    p_gen.add_argument(&quot;--prompt&quot;, required=True)
    p_gen.add_argument(&quot;--num-new-tokens&quot;, type=int, default=500)
    p_gen.add_argument(&quot;--temperature&quot;, type=float, default=0.8)
    p_gen.add_argument(&quot;--top-k&quot;, type=int, default=40)
    p_gen.add_argument(&quot;--seed&quot;, type=int, default=None)
    p_gen.set_defaults(func=generate)

    args = parser.parse_args()
    args.func(args)   # dispatch to train() or generate()


if __name__ == &quot;__main__&quot;:
    main()

</pre></div>


<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=39026" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/06/14/back-to-basics-build-your-own-llm-from-scratch/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">39026</post-id>	</item>
		<item>
		<title>Exploring Epicure the Food Embedding Model</title>
		<link>https://thejeshgn.com/2026/06/11/exploring-epicure-the-food-embedding-model/</link>
					<comments>https://thejeshgn.com/2026/06/11/exploring-epicure-the-food-embedding-model/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Thu, 11 Jun 2026 09:54:38 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[arXiv]]></category>
		<category><![CDATA[Embedding Models]]></category>
		<category><![CDATA[Free and Open Source]]></category>
		<category><![CDATA[Metapath2Vec]]></category>
		<category><![CDATA[Word2Vec]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=38996</guid>

					<description><![CDATA[FlavorGraph is a large-scale graph network that combines data from over a million recipes with chemical compound information from 1,500+ flavor molecules to predict ingredient pairings. It uses graph embedding methods to represent foods as dense vectors, enabling data-driven food pairing suggestions that go beyond human or chef intuition. In FlavorGraph, the chemical&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p><a href="https://www.nature.com/articles/s41598-020-79422-8" target="_blank" rel="noreferrer noopener">FlavorGraph</a> is a large-scale graph network that combines data from over a million recipes with chemical compound information from 1,500+ flavor molecules to predict ingredient pairings. It uses graph embedding methods to <a href="https://github.com/lamypark/FlavorGraph" target="_blank" rel="noreferrer noopener">represent foods as dense vectors</a>, enabling data-driven food pairing suggestions that go beyond human or chef intuition. In FlavorGraph, the chemical and recipe context signals are fused at training time via a fixed metapath design, leaving no inference-time knob to adjust their relative weights in the final embeddings.</p>



<p>One could call <a href="https://arxiv.org/abs/2605.22391" target="_blank" rel="noreferrer noopener">Epicure</a> an enhanced FlavorGraph. It builds on <a href="https://github.com/lamypark/FlavorGraph" target="_blank" rel="noreferrer noopener">FlavorGraph</a> to produce 300-D embeddings, but instead of a single embedding that combines both chemical and recipe context signals. It has three embedding models <a href="https://huggingface.co/Kaikaku/epicure-cooc" target="_blank" rel="noreferrer noopener">Cooc</a>, <a href="https://huggingface.co/Kaikaku/epicure-chem" target="_blank" rel="noreferrer noopener">Chem</a>, and <a href="https://huggingface.co/Kaikaku/epicure-core" target="_blank" rel="noreferrer noopener">Core</a>. That way, as a user, you can choose the embedding you want. It also includes more recipes from other languages, not just English. Recipes from English, Chinese, Russian, Vietnamese, Spanish, Turkish, Indonesian, German, and Indian-English are included. They are machine-translated into English for this. It&#8217;s good to see Indian recipes covered as well.</p>



<p>They have also normalized the raw ingredient strings to <a target="_blank" href="https://huggingface.co/Kaikaku/epicure-core/blob/main/itos.json" rel="noreferrer noopener">1,790 canonical entries</a> using an LLM. I found this interesting. I also found that this doesn&#8217;t necessarily cover everything that an Indian recipe needs. For example, there are only coriander and coriander_root. It&#8217;s usually coriander seeds or coriander leaves in India. In their list, there is no way to differentiate it.</p>



<p>The most interesting part is how they traverse the graph to construct metapaths for training Metapath2Vec models. The three models follow different logic to construct it. But they all have the same architecture and hyperparameters: <code>300-dim embeddings, walks_per_node=100, walk_length=50, context_size=7, 5 negative samples, batch_size=32,768, lr=0.0025, 20 epochs, no warm restart</code>.</p>



<p>From the <a href="https://arxiv.org/html/2605.22391v1#S2" target="_blank" rel="noreferrer noopener">paper</a>, the methodology (I have embedded the <a href="https://thejeshgn.com/wp-content/uploads/2026/06/epicure_method.svg">SVG flowchart </a>below) is easy to understand and, if required, can be replicated. But the raw data is not available, though sources are mentioned. Maybe it can be replicated with a different dataset. They have also used LLMs quite a bit in the data and training pipeline. For me, constructing the metapaths was the most interesting part.</p>



<ol class="wp-block-list">
<li>Epicure-Cooc. Walks the Cooc graph: pure I–I random walks weighted by NPMI. No compound nodes.</li>



<li>Epicure-Core. Walks the typed-compound graph and injects pure I–I walks at <code>--ii_repeat=10</code> alongside the typed-compound metapaths. Edge transitions are weighted so I–C hops are not oversampled relative to the smaller I–I edge set. The resulting embedding blends chemical and recipe-context signal.</li>



<li>Epicure-Chem. Walks the typed-compound graph but with &#8211;ii_repeat=0: the I–I templates are absent and the only walks the skip-gram sees are compound-mediated. The chemistry extreme of the family.</li>
</ol>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><a href="https://thejeshgn.com/wp-content/uploads/2026/06/epicure_method.svg"><img decoding="async" src="https://thejeshgn.com/wp-content/uploads/2026/06/epicure_method.svg" alt="" class="wp-image-38997" style="width:500px"/></a><figcaption class="wp-element-caption">Epicure Methodology Flowchart</figcaption></figure>
</div>


<p>You can explore the <a href="https://huggingface.co/spaces/Kaikaku/epicure-explorer" target="_blank" rel="noreferrer noopener">embedding online</a> or using a simple script below. In my exploration, I found that it is okay to use the nearest neighbors for an ingredient in a recipe, in chemistry, or in mixed contexts. But I don&#8217;t think it&#8217;s at a level where we could just replace ingredients. It also doesn&#8217;t have any information about allergens, texture, etc. But it is small and can be used to build on top of it.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><a href="https://huggingface.co/spaces/Kaikaku/epicure-explorer"><img loading="lazy" decoding="async" width="1353" height="1102" src="https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer.png" alt="Online explorere Epicure three sibling ingredient embeddings." class="wp-image-39004" srcset="https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer.png 1353w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-300x244.png 300w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-1024x834.png 1024w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-768x626.png 768w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-720x586.png 720w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-520x424.png 520w, https://thejeshgn.com/wp-content/uploads/2007/07/epicure_embeddings_explorer-320x261.png 320w" sizes="auto, (max-width: 1353px) 100vw, 1353px" /></a><figcaption class="wp-element-caption"><a href="https://huggingface.co/spaces/Kaikaku/epicure-explorer">Online explorer</a> Epicure three sibling ingredient embeddings.</figcaption></figure>
</div>

<div class="wp-block-syntaxhighlighter-code "><pre class="brush: python; title: ; notranslate">
# /// script
# requires-python = &quot;&gt;=3.11&quot;
# dependencies = &#x5B;
#   &quot;numpy&quot;,
#   &quot;huggingface_hub&quot;,
# ]
# ///

from epicure import Epicure

m = Epicure.from_pretrained(&quot;Kaikaku/epicure-core&quot;)

print(&quot;neighbors : chicken&quot;)
print(m.neighbors(&quot;chicken&quot;, k=5))

print(&quot;neighbors : coriander&quot;)
print(m.neighbors(&quot;coriander&quot;, k=5))

print(&quot;slerp : rice, cuisine:South_Asian&quot;)
print(m.slerp(&quot;rice&quot;, &quot;cuisine:South_Asian&quot;, theta_deg=30, k=5))

print(&quot;closest_mode : chocolate&quot;)
print(m.closest_mode(&quot;chocolate&quot;, kind=&quot;factor&quot;, k=3))
</pre></div>


<p></p>



<p class="has-luminous-vivid-amber-background-color has-background">To run the above script, include the epicure.py file from the <a href="https://huggingface.co/Kaikaku/epicure-core/tree/main" target="_blank" rel="noreferrer noopener">HF repo</a> in the same folder. The epicure package on PyPI is a different package.</p>



<p></p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: bash; title: ; notranslate">
# Output of above script
neighbors : chicken
&#x5B;(&#039;pork&#039;, 0.5807677507400513), (&#039;beef&#039;, 0.5712239742279053), (&#039;chicken_broth&#039;, 0.5498523116111755), (&#039;peanut&#039;, 0.5233361124992371), (&#039;cream_of_chicken_soup&#039;, 0.5217206478118896)]
neighbors : coriander
&#x5B;(&#039;cumin&#039;, 0.7044016122817993), (&#039;scallion&#039;, 0.6711025238037109), (&#039;chili_pepper&#039;, 0.6609357595443726), (&#039;turmeric&#039;, 0.6468461155891418), (&#039;chicken_broth&#039;, 0.6463753581047058)]
slerp : rice, cuisine:South_Asian
&#x5B;(&#039;turmeric&#039;, 0.7607102225586673), (&#039;mustard_seed&#039;, 0.756659547118539), (&#039;fenugreek_seed&#039;, 0.7468295496269882), (&#039;coriander&#039;, 0.7428115853809022), (&#039;cumin&#039;, 0.7388300342030064)]
closest_mode : chocolate
&#x5B;(&#039;F_15/M0&#039;, &#039;American sweet confections and dessert bases&#039;, 0.7751955986022949), (&#039;F_7/M1&#039;, &#039;Sweet liqueurs and confections&#039;, 0.7553147673606873), (&#039;F_8/M4&#039;, &#039;Sweet confections and dessert ingredients&#039;, 0.7396374344825745)]

</pre></div>


<h3 class="wp-block-heading">Definitions</h3>



<p><a href="https://en.wikipedia.org/wiki/Word2vec" target="_blank" rel="noreferrer noopener">Word2Vec</a> is a neural method for building word embeddings from raw text. Words appearing in similar contexts end up close together in the vector space. It comes in two variants skip-gram and CBOW.</p>



<ul class="wp-block-list">
<li>skip-gram: predict the surrounding context words from a target word</li>



<li>CBOW: predict the target word from its surrounding context</li>
</ul>



<p><a href="https://ericdongyx.github.io/papers/KDD17-dong-chawla-swami-metapath2vec.pdf" target="_blank" rel="noreferrer noopener">Metapath2Vec (PDF)</a> It does random walks through the graph guided by a defined metapath, then feeds those walks to a skip-gram model to learn node embeddings. It works well for Heterogeneous graphs. Basically, it&#8217;s like creating sentences based on the metapath, then running Word2Vec skip-gram on them.</p>



<ul class="wp-block-list">
<li>A metapath is a typed node sequence, such as author – paper – venue – paper – author. It&#8217;s predefined by manually by the model designer.</li>
</ul>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=38996" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/06/11/exploring-epicure-the-food-embedding-model/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">38996</post-id>	</item>
		<item>
		<title>Embedding user code in your app using Extism</title>
		<link>https://thejeshgn.com/2026/05/19/embedding-user-code-in-your-app-using-extism/</link>
					<comments>https://thejeshgn.com/2026/05/19/embedding-user-code-in-your-app-using-extism/#respond</comments>
		
		<dc:creator><![CDATA[Thejesh GN]]></dc:creator>
		<pubDate>Tue, 19 May 2026 14:40:06 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[Free and Open Source]]></category>
		<category><![CDATA[Power User]]></category>
		<category><![CDATA[WASM]]></category>
		<category><![CDATA[Web Assembly]]></category>
		<guid isPermaLink="false">https://thejeshgn.com/?p=38848</guid>

					<description><![CDATA[Every application I love has some kind of power-user mode where I can add my own code or scripting to make it useful to me. Simple examples are Firefox with its addons or VLC with its plugins. Ideally, any significant or valuable application should have such a feature. But as a developer or&#46;&#46;&#46;]]></description>
										<content:encoded><![CDATA[
<p>Every application I love has some kind of power-user mode where I can add my own code or scripting to make it useful to me. Simple examples are Firefox with its addons or VLC with its <a target="_blank" href="https://addons.videolan.org/browse?cat=325&amp;ord=latest" rel="noreferrer noopener">plugins</a>. Ideally, any significant or valuable application should have such a feature.</p>



<p>But as a developer or builder, it&#8217;s not easy to build such a feature securely. I have tried it before with embedding Lua with some restrictions. But it isn&#8217;t great in sandboxing. I&#8217;ve always had an eye on the WASM-based implementation because browsers have decent sandboxing. Very recently, when I was <a href="https://thejeshgn.com/2026/05/08/weekly-notes-19-2026/" target="_blank" rel="noreferrer noopener">updating Navidrome</a>, I came across its plugin system, which uses <a href="https://extism.org/" target="_blank" rel="noreferrer noopener">Extism</a> and is based on WebAssembly. Ideally, at some point, even WordPress can do this, allowing us to sandbox plugins. </p>



<p>So I wrote a simple plugin to try out this infrastructure. I wrote a plugin with internet access to a given URL. Its only goal is to test the plugin&#8217;s ability to make web API calls in a controlled way. And  to explore how plugins work. The language I chose for the plugin is Python, but there are many <a href="https://extism.org/docs/concepts/pdk" target="_blank" rel="noreferrer noopener">better supported languages</a>. I tested it using CLI. And it&#8217;s not difficult to call it from a host language.</p>



<p>I installed and used Extism <a href="https://extism.org/docs/install" target="_blank" rel="noreferrer noopener">CLI</a>, version 1.6.2. Direct binary download from the repository. I wrote a simple Python script that, as a plugin, expects a JSON input containing URL and payload for a web post. There are other ways to pass the data to the plugin, but JSON seemed most flexible.</p>



<p>Inside the plugin, I read the JSON input, parse the parameters, and make a web POST request using <code>Existims Http</code>. Remember, you need to use the <a href="https://github.com/extism/python-pdk" target="_blank" rel="noreferrer noopener">Python PDK (Plugin Development Kit)</a>. Those PDKs can have some limitations of their own. So be mindful when choosing a plugin language. </p>



<p>Also as you can see, a plugin can have many functions. The ones annotated with <code>@extism.plugin_fn</code> can be called from the host environment. The code is easy to understand.</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: python; title: ; notranslate">
#webhook.py
import json
import extism
from extism import Http

@extism.plugin_fn
def post_json():
    params = extism.input_json()

    url = params&#x5B;&quot;url&quot;]
    payload = params&#x5B;&quot;payload&quot;]

    response = Http.request(
        url,
        meth=&quot;POST&quot;,
        headers={&quot;Content-Type&quot;: &quot;application/json&quot;},
        body=json.dumps(payload),
    )

    result = {
        &quot;status&quot;: response.status_code,
        &quot;body&quot;: response.data_str(),
    }

    extism.output_str(json.dumps(result))
</pre></div>


<p>To call this plugin first I need to build it. To build use the <code>extism-py</code> which is installed as part of Python PDK. Then run</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: bash; title: ; notranslate">
extism-py webhook.py webhook.wasm
</pre></div>


<p>That will produce <code>webhook.wasm</code>, the plugin code that one would probably ship. You can call or invoke it via the CLI for testing. As you can see, I have to grant internet access to the WASM by allowing access to the url <code>webhook.site</code>.</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: bash; title: ; notranslate">
extism call webhook.wasm post_json \
--allow-host &quot;webhook.site&quot; --wasi \
--input &#039;{ &quot;url&quot;: &quot;https://webhook.site/0b757518-7120-4919-a12f-252d3dfbc8b5&quot;, &quot;payload&quot;: { &quot;name&quot;: &quot;Thejesh from inside the sandbox&quot;, &quot;seq&quot;: 1 } }&#039;
</pre></div>


<p>But in real world you would be calling it from host code/program. Let&#8217;s say your host language is also Python, then you can call this WASM plugin using the following piece of code. Again, it&#8217;s not that difficult.</p>


<div class="wp-block-syntaxhighlighter-code "><pre class="brush: plain; title: ; notranslate">
# host.py
# /// script
# dependencies = &#x5B;&quot;extism&quot;]
# ///

import json
import extism

wasm_file = &quot;webhook.wasm&quot;

input_data = {
    &quot;url&quot;: &quot;https://webhook.site/dc895b92-0b21-46db-a9c0-766dd87e8b0f&quot;,
    &quot;payload&quot;: {
        &quot;name&quot;: &quot;Thejesh from inside the sandbox&quot;,
        &quot;seq&quot;: 1,
    },
}

manifest = {
    &quot;wasm&quot;: &#x5B;{&quot;path&quot;: wasm_file}],
    &quot;allowed_hosts&quot;: &#x5B;&quot;webhook.site&quot;],
}

with extism.Plugin(
    manifest,
    wasi=True,
) as plugin:
    result = plugin.call(&quot;post_json&quot;, json.dumps(input_data).encode())
    print(result.decode())

</pre></div>


<p>All of a sudden, safe plugin implementation doesn&#8217;t sound that difficult, isn&#8217;t it?</p>



<hr class="wp-block-separator has-text-color has-vivid-purple-color has-alpha-channel-opacity has-vivid-purple-background-color has-background is-style-dots"/>



<div class="wp-block-columns has-background is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex" style="background:linear-gradient(137deg,rgb(255,206,236) 0%,rgb(152,150,240) 62%)">
<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">
<p></p>



<p>You can read this blog using <a href="https://feeds.thejeshgn.com/thejeshgn" target="_blank" rel="noreferrer noopener">RSS Feed</a>. But if you are the person who loves getting emails, then you can join my readers by <a href="https://thejeshgn.com/subscribe/" target="_blank" rel="noreferrer noopener">signing</a> up.</p>


<div class="wp-block-jetpack-subscriptions__supports-newline wp-block-jetpack-subscriptions__show-subs wp-block-jetpack-subscriptions">
		<div>
			<div>
				<div>
					<p style="width: 25%;max-width: 100%;">
						<a href="https://thejeshgn.com/?post_type=post&#038;p=38848" style="width: calc(100% - 10px);font-size: 16px;padding: 15px 23px 15px 23px;margin: 0; margin-left: 10px;border-color: black;border-radius: 6px;border-width: 2px; background-color: #000000; color: #FFFFFF; text-decoration: none; white-space: nowrap; margin-left: 0">Subscribe</a>
					</p>
				</div>
			</div>
		</div>
	</div></div>



<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow" style="flex-basis:5px"></div>
</div>



<p> </p>
]]></content:encoded>
					
					<wfw:commentRss>https://thejeshgn.com/2026/05/19/embedding-user-code-in-your-app-using-extism/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">38848</post-id>	</item>
	</channel>
</rss>
