# How to iterate all pages in custom folders in jekyll

**URL:** <https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311>\
**Category:** Help\
**Created:** [May 6, 2022, 5:04am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311 "2022-05-06T05:04:08Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 6, 2022, 5:04am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/1 "2022-05-06T05:04:08Z")

</div>

HI Team. I have implemented one githubpages website.Andmy project structure is

repository root folder  
Docs folder  
test folder  
file.md  
inner-test folder  
file1.md  
file2.md  
file3.md  
sample folder  
file4.md  
file5.md  
innersample folder  
file6.md  
test2 folder   
file7.md  
home.md  
index.md  
\_config.yml

```
	So now I want to iterate through all the folders and get the page.name and page.content and page.url to implement the search functionality can you pls help me.
	
	I tried below code and not working. 

```

{% for post in site.docs %}  
“{{ post.url | slugify }}”: {  
“title”: “{{ post.name | xml\_escape }}”,  
“category”: “{{ post.category | xml\_escape }}”,  
“content”: {{ post.content | strip\_html | strip\_newlines | jsonify }},  
“url”: “{{ post.url | xml\_escape }}”  
}  
{% unless forloop.last %},{% endunless %}  
{% endfor %}

I provided my test repo URL also here. pls help.  
[testrepo/docs at main · vanamsandeep/testrepo (github.com)](https://github.com/vanamsandeep/testrepo/tree/main/docs)

---

<div class="post-metadata">

**Author:** ![BillRaymond](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/billraymond/32/5963_2.png) [@BillRaymond](https://talk.jekyllrb.com/u/BillRaymond)\
**Post date:** [May 6, 2022, 6:51pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/2 "2022-05-06T18:51:37Z")

</div>

There are two primary document collections in Jekyll: posts and pages. Another will be [collections](https://jekyllrb.com/docs/collections/) if you use them.

Pages and posts may contain their own front matter, so I suggest you treat each one separately.

## ✍ List of all posts

Read the Jekyll array containing a list of all the posts:

```auto
{% assign posts = site.posts %}
{% for post in posts %}
    title: {{post.title}}
{% endfor %}

```

## 📄 List of all pages

Read the Jekyll array containing a list of all the pages:

```auto
{% assign pages = site.pages | where_exp: 'page', 'page.title' %}
{% for page in pages %}
    title: {{page.title}}
{% endfor %}

```

Notice I used a `where_exp` filter in the list of pages. Sometimes a page does not have a `title` in the front matter. That will remove untitled pages from the listing. If you want all of them, even without a title, then the code would look like this instead:

```auto
{% assign pages = site.pages %}
{% for page in pages %}
    title: {{page.title}}
{% endfor %}

```

## ➡ More options

I recommend you check out the list of built-in Jekyll site variables. You can learn about the different arrays you can access that Jekyll automatically collects for you. There may be other document types (like HTML docs) that you want to include as well.

> **[Variables](https://jekyllrb.com/docs/variables/#site-variables)**
>
> Jekyll traverses your site looking for files to process. Any files with front matter are subject to processing. For each of these files, Jekyll makes a variety of data available via Liquid. The following is a reference of the available data.

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 7, 2022, 2:49am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/3 "2022-05-07T02:49:09Z")

</div>

Thank you it worked. I have one more issue. How can I remove the below details from page.content: by using filters in above loop  
I have \<img src fields in content how can I remove using filters.  
I tried “content”: {{ page.content | markdownify | strip\_html | markdownify | jsonify }}, but not working.  
Also tried : “content”: {{ page.content | strip\_html | strip\_newlines | jsonify }}  
tried multiple remove and replace filters but none of them working.

Screen shot is:  
 ![image](https://canada1.discourse-cdn.com/flex030/uploads/jekyllrb/original/2X/7/71fed11c8ac60790ad6e6b31bd2f18036443335a.png)

---

<div class="post-metadata">

**Author:** ![BillRaymond](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/billraymond/32/5963_2.png) [@BillRaymond](https://talk.jekyllrb.com/u/BillRaymond)\
**Post date:** [May 7, 2022, 7:21pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/4 "2022-05-07T19:21:24Z")

</div>

Please share the original content you are working with.

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 7, 2022, 7:24pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/5 "2022-05-07T19:24:55Z")

</div>

I have 100’s of markdown files and one of the sample markdown file is

```auto
---
permalink: "/"
---

# SEARCH :mag: 
You can <b>[Search](ww.example.com)</b> with a combination of keywords. For documentation results, there is always a Markdown file listed (*.md).

# Test

```

* * *

Html code for iterating:

```auto
---
layout: search
---
<form action="/search.html" method="get">
    <label for="search-box">Search</label>
    <input type="text" id="search-box" name="query">
    <input type="submit" value="search">
</form>

<ul id="search-results"></ul>

 window.store = {
    {% for page in site.pages %}
      "{{ page.url | slugify }}": {
        "title": "{{ page.name | xml_escape }}",
        "content": "{{ page.content | strip_html | strip_newlines | replace: '"', '\"' }}",
        "url": "{{ page.url | xml_escape }}"
      }
      {% unless forloop.last %},{% endunless %}
    {% endfor %}

  };

<script src="/js/lunr.min.js"></script>
<script src="/js/search.js"></script>

```

---

<div class="post-metadata">

**Author:** ![rdyar](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/rdyar/32/710_2.png) [@rdyar](https://talk.jekyllrb.com/u/rdyar)\
**Post date:** [May 7, 2022, 10:11pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/6 "2022-05-07T22:11:16Z")

</div>

I tried to reformat your post, I think it is correct?

when posting code you should use the \</\> code button in the editor or wrap it in 2 sets of 3 back tics ( ```)

---

<div class="post-metadata">

**Author:** ![rdyar](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/rdyar/32/710_2.png) [@rdyar](https://talk.jekyllrb.com/u/rdyar)\
**Post date:** [May 7, 2022, 10:17pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/7 "2022-05-07T22:17:19Z")

</div>

> [@sandeepv](#):
>
> Also tried : “content”: {{ page.content | strip\_html | strip\_newlines | jsonify }}

when you do that do you get what is in your screenshot? if no what do you get?

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 8, 2022, 2:58am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/8 "2022-05-08T02:58:03Z")

</div>

Yes the this is correct. I’m implementing the search functionality in my blog by using lunr.js search functionality. For that I need to iterate all the pages in the blog using jekyll liquid tag to form as json structure. So while using the above code json structure is breaking because of some html tags. I need help on removing those html tags and make it as valid json structure.  
My Markdown file code is:

```auto
# SEARCH :mag: 
You can <b>[Search](https://example/search?search=1234)</b> with a combination of keywords for any of our for documentation results, there is always a Markdown file listed (*.md).

support requests to [examplesupport@example.com](mailto:example.example.com)

```

Below the search.html code to iterate and get it as json file

```auto
---
layout: search
---
<form action="/testrepo/search.html" method="get">
    <label for="search-box">Search</label>
    <input type="text" id="search-box" name="query">
    <input type="submit" value="search">
</form>

<ul id="search-results"></ul>

<script>
 window.store = {
    {% for page in site.pages %}
      "{{ page.url | slugify }}": {
        "title": "{{ page.name | xml_escape }}",
        "content": "{{ page.content | strip_html | strip_newlines }}",
        "url": "{{ page.url | xml_escape }}"
      }
      {% unless forloop.last %},{% endunless %}
    {% endfor %}

  };
</script>
<script src=/js/lunr.min.js"></script>
<script src="/js/search.js"></script>

```

With this code it’s iterating but it’s breaking like below.

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/jekyllrb/original/2X/8/8bd2a90f64491fe5e1062aaab09f9e6bbce152e9.png)

I’m also pasting my code repository and URL:  
Github repo: [vanamsandeep/testrepo (github.com)](https://github.com/vanamsandeep/testrepo)  
URL: [Jekyll Actions Demo with leap-day theme](https://vanamsandeep.github.io/testrepo/search)

pls suggest me on this

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 8, 2022, 3:00am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/9 "2022-05-08T03:00:06Z")

</div>

I’m implementing the search functionality in my blog by using lunr.js search functionality. For that I need to iterate all the pages in the blog using jekyll liquid tag to form as json structure. So while using the above code json structure is breaking because of some html tags. I need help on removing those html tags and make it as valid json structure.  
My Markdown file code is:

```auto
# SEARCH :mag: 
You can <b>[Search](https://example/search?search=1234)</b> with a combination of keywords for any of our for documentation results, there is always a Markdown file listed (*.md).

support requests to [examplesupport@example.com](mailto:example.example.com)

```

Below the search.html code to iterate and get it as json file

```auto
---
layout: search
---
<form action="/testrepo/search.html" method="get">
    <label for="search-box">Search</label>
    <input type="text" id="search-box" name="query">
    <input type="submit" value="search">
</form>

<ul id="search-results"></ul>

<script>
 window.store = {
    {% for page in site.pages %}
      "{{ page.url | slugify }}": {
        "title": "{{ page.name | xml_escape }}",
        "content": "{{ page.content | strip_html | strip_newlines }}",
        "url": "{{ page.url | xml_escape }}"
      }
      {% unless forloop.last %},{% endunless %}
    {% endfor %}

  };
</script>
<script src=/js/lunr.min.js"></script>
<script src="/js/search.js"></script>

```

With this code it’s iterating but it’s breaking like below.

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/jekyllrb/original/2X/8/8bd2a90f64491fe5e1062aaab09f9e6bbce152e9.png)

I’m also pasting my code repository and URL:  
Github repo: [vanamsandeep/testrepo (github.com)](https://github.com/vanamsandeep/testrepo)  
URL: [Jekyll Actions Demo with leap-day theme](https://vanamsandeep.github.io/testrepo/search)

pls suggest me on this

---

<div class="post-metadata">

**Author:** ![BillRaymond](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/billraymond/32/5963_2.png) [@BillRaymond](https://talk.jekyllrb.com/u/BillRaymond)\
**Post date:** [May 9, 2022, 5:59pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/10 "2022-05-09T17:59:39Z")

</div>

I went to your repo and did not see the code I shared on there (or a variant of it). Are you testing on a specific branch?

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 9, 2022, 6:00pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/11 "2022-05-09T18:00:42Z")

</div>

You can edit on main branch this is my testrepo only

---

<div class="post-metadata">

**Author:** ![sandeepv](https://avatars.discourse-cdn.com/v4/letter/s/7cd45c/32.png) [@sandeepv](https://talk.jekyllrb.com/u/sandeepv)\
**Post date:** [May 9, 2022, 6:02pm UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/12 "2022-05-09T18:02:10Z")

</div>

> <https://github.com/vanamsandeep/testrepo/blob/main/docs/search.html>

---

<div class="post-metadata">

**Author:** ![KhBh](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.jekyllrb.com/khbh/32/4845_2.png) [@KhBh](https://talk.jekyllrb.com/u/KhBh)\
**Post date:** [March 2, 2023, 1:21am UTC](https://talk.jekyllrb.com/t/how-to-iterate-all-pages-in-custom-folders-in-jekyll/7311/13 "2023-03-02T01:21:05Z")

</div>

In case it helps anyone searching this later, here is my working Lunr.js search index:

> <https://github.com/buddhist-uni/buddhist-uni.github.io/blob/deb554ca248c3c5bb6bf0de9ee496ba399aa3a9b/assets/js/search_index.js>

The trickiest part was filtering out liquid code that sometimes appears in `page.content`
