# Replacing a slow include with a custom Ruby Tag

**URL:** <https://talk.jekyllrb.com/t/replacing-a-slow-include-with-a-custom-ruby-tag/6064>\
**Category:** Share\
**Created:** [May 21, 2021, 2:58pm UTC](https://talk.jekyllrb.com/t/replacing-a-slow-include-with-a-custom-ruby-tag/6064 "2021-05-21T14:58:10Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![KhBh](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/khbh/32/4845_2.png) [@KhBh](https://talk.jekyllrb.com/u/KhBh)\
**Post date:** [May 21, 2021, 2:58pm UTC](https://talk.jekyllrb.com/t/replacing-a-slow-include-with-a-custom-ruby-tag/6064/1 "2021-05-21T14:58:10Z")

</div>

On [most pages](https://buddhistuniversity.net/content/monographs/miracle-of-mindfulness_tnh) on [my website](https://buddhistuniversity.net/), I have a footer which recommends similar items in the archive. Since I have very little text _about_ these items, and lots of rich metadata, I decided to implement my own content similarity engine and, since I was using GitHub Pages, I did this [entirely in Liquid](https://github.com/buddhist-uni/buddhist-uni.github.io/blob/v1.4/_includes/similar_content_footer.html).

Fast forward a few months, and the site now has [over a thousand](https://buddhistuniversity.net/content/) pieces of content. Since the similarity algo is O(n^2), this blew up and the other day my build times exceeded the GitHub Pages timeout of 10 mins / build.

In case it’s useful for anyone else, here is what I did to fix it:

1. I switched my site from the default GitHub Pages deploy to [a custom GitHub workflow](https://github.com/buddhist-uni/buddhist-uni.github.io/blob/master/.github/workflows/build.yaml). A few gotchyas there. Look at the linked workflow file for what I eventually landed on. Not only does this overcome the 10 min build timeout, but it also (finally!) let me upgrade Jekyll and use (make) custom plugins:
2. This overcame the timeout problem, but the increased overhead actually increased my build times. I upgraded to Jekyll 4.2, and changed all my `site.foos | where: "slug", bar | first` with the more sensible `site.foos | find: "slug", bar`. This did make my build _slightly_ faster, but I was honestly underwhelmed by the improvement, so then…
3. I [rewrote](https://github.com/buddhist-uni/buddhist-uni.github.io/commit/606d0461e3433d5f8fa121a9e3e16cf721afb685) the [offending `_include`](https://github.com/buddhist-uni/buddhist-uni.github.io/blob/v1.4/_includes/similar_content_footer.html) code [in Ruby as a custom Tag `_plugin`](https://github.com/buddhist-uni/buddhist-uni.github.io/blob/master/_plugins/similar_content.rb). This was relatively straightforward, but came with a few gotchyas:
  1. My original include [called another include](https://github.com/buddhist-uni/buddhist-uni.github.io/blob/v1.4/_includes/similar_content_footer.html#L87). In order to render a vanilla include from Ruby-land, I had to copy-paste (!) [the IncludeTag’s `load_cached_partial` function](https://github.com/jekyll/jekyll/blob/v4.2.0/lib/jekyll/tags/include.rb#L140) into my Tag (😬) to be able to load the include I wanted to render. (Perhaps that Jekyll method could be made public / static?)
  2. Sometimes when you fetch a page object from `context` it comes back as a `Document` and sometimes it comes back as a `DocumentDrop`? Eventually, I figured out that the best way to get data off these things consistently and reliably was to call `.to_liquid.to_h` on any such object first, converting it to a plain old Ruby Hash.
  3. Lastly, this hacking and debugging was quite slow at first, as I had to stop and restart the server each time to reload the Ruby file (unlike Liquid changes which can reload on the fly). Eventually I figured out a hack 🙈 A custom tag can accept arbitrary `text` after its name. So, I passed that `@text` into an `eval()` in my render function, thus allowing me to run arbitrary code from the proper context without having to restart the server. Not sure if there is a nicer way to drop a debugger in, but this worked for me 😈

At the end of the day, I managed to speed up this particular O(n^2) path by about 10x and cut my overall build time in about half: a result I’m quite happy with! If you have any thoughts or feedback, please comment below, and thanks for reading! 😄

---

<div class="post-metadata">

**Author:** ![KhBh](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/khbh/32/4845_2.png) [@KhBh](https://talk.jekyllrb.com/u/KhBh)\
**Post date:** [December 6, 2022, 7:18am UTC](https://talk.jekyllrb.com/t/replacing-a-slow-include-with-a-custom-ruby-tag/6064/2 "2022-12-06T07:18:17Z")

</div>

Fast-forward a year and a half and my site has grown large enough that even this nativized O(n²) algorithm is too slow.

Thankfully, k-NN is a known algorithm, and by precomputing some binary space partitions you can narrow down the search space dramatically.

In my case, that meant simply keeping around N sets of all the posts with a given tag, and then merging those sets to find the posts with the highest tag overlap.

My build is now ~7x faster 😈 Happy Jekylling!
