Documentation
Contents
- Adding a site to the public search
- Adding a site to the Website Search Tool
- What benefits do I get from a Full listing?
- Does my site have to appear in the public search if I want to use the Website Search Tool?
- Can I change a Basic listing to a Full listing?
- What happens when my Full listing expires?
- How do I verify ownership of my site?
- Do I have to complete a Full listing in one session?
- Managing my site
- How do I add a search box to my site?
- What is the specification for the API?
- How does the indexing process work?
- How frequently are sites reindexed?
- How does the relevancy tuning work?
- Using the search
- Support
Adding a site to the public search
What types of site can be submitted to the public search?
Users are encouraged to submit any personal websites which they believe will improve the public search, not just their own sites. See the Submission Guidelines section in the Terms of Use for further details, including a definition of personal websites.
How do I add a site to the public search?
Use the Add Site link at the top. Only people with a Full listing can suggest a site for a Basic listing, and it has to be approved by a moderator before it is indexed. Note that only site owners can submit a Free Trial or Full listing, because the submission process requires verifying ownership of the site.
Adding a site to the Website Search Tool
What benefits do I get from a Full listing?
The Full listing provides a number of benefits, for example more frequent indexing, indexing a higher number of pages, access to Manage Site to configure indexing and trigger on-demand reindexing, enabling the API, etc. See Website Search Tool for further details.
Does my site have to appear in the public search if I want to use the Website Search Tool?
No. When listing your site one of the questions is "Include in public search", to which you can select No. This allows you to use the Website Search Tool even if your site isn't a personal or independent website. Note that, as per the Terms of Use, a moderator may exclude a Full listing from the public search.
Can I change a Basic listing to a Full listing?
Yes. Just resubmit the site, and select Full for the Listing tier.
What happens when my Full listing expires?
When a Full listing expires, it will revert to a Free Trial listing for one month. To convert a Free Trial to a Full listing, simply log in to Manage Site, select the Subscriptions tab, and click Purchase. After the Free Trial expires it will revert to a Basic listing. To convert a Basic to a Full listing, you will need to renew via Add Site, although you will not need to verify ownership of your site again. If you have added a search box to your site with the one line of code provided by the Website Search Tool, it will continue to work whatever the listing type, although the behaviour will change as described below.
How do I verify ownership of my site?
The process similar to that which you may have used for other services, i.e. you upload a specific piece of content to the domain's root or TXT record. The Add Site link will guide you through the process.
Do I have to complete a Full listing in one session?
No. For a Full listing, if you are unable to complete the process in one session, you can resubmit at a later time to pick up where you left off.
Managing my site
How do I add a search box to my site?
The simplest way is to add a single line of code to a page of your site, normally a dedicated search page such as /search/, e.g. (for michael-lewis.com):
<script src="https://searchmysite.net/static/js/sms-search.js" data-domain="michael-lewis.com"></script>
Set data-domain to your domain. The script works whatever your listing tier is (Basic, Free Trial, or Full), but the behaviour differs:
- With a Full listing (or Free Trial) it shows a search box and the results on your own site using the API.
- With a Basic listing it shows a search box on your own site but the results on the searchmysite.net results page.
Pages using a search box added in this way can be deep-linked using the q parameter, e.g. https://michael-lewis.com/search/?q=antarctica.
Note that, for static site generators which escape HTML in markdown
(e.g. Hugo), you will need to allow raw HTML (e.g. add
[markup.goldmark.renderer] unsafe = true to your config).
Alternatively, you can use the API directly to build your own custom search experience, either server-side or client-side on your own domain. If you prefer to do it manually without the API, a simple option is to have a form which takes a query and a domain hidden parameter containing the value of the domain to which to restrict results, e.g. (for michael-lewis.com):
<form action="https://searchmysite.net/search/">
<input type="search" name="q" ></input>
<input type="hidden" name="domain" value="michael-lewis.com"></input>
<input type="submit" value="Search"></input>
</form>
What is the specification for the API?
In summary, queries take the form /api/v1/search/<domain>?q=*, where parameters are:
- <domain>: the domain being searched (mandatory)
- q: query string (mandatory)
- page: the page number from which multi-page results should start (optional, default 1)
- resultsperpage: the number of results per page (optional, default 10)
Results are returned in the following format, with all fields optional apart from id and url:
{
"params": {
"q": "*",
"page": 1,
"resultsperpage": 10,
}
"totalresults": 40,
"results": [
{
"id": "https://server/path",
"url": "https://server/path",
"title": "Page title",
"author": "Author",
"description": "Page description",
"tags": ["tag1", "tag2"],
"page_type": "Page type, e.g. article",
"page_last_modified": "2020-07-17T00:00:00+00:00",
"published_date": "2020-07-17T00:00:00+00:00",
"language": "en",
"indexed_inlinks": ["inlink1", "inlink2"],
"indexed_outlinks": ["outlink1", "outlink2"],
"fragment": ["text before the search ", "query", " and text after"]
}
]
}
How does the indexing process work?
The indexing process first checks a robots.txt, and will obey any rules there. If the robots.txt allows, it will then load the home page and web feed (if a web feed was configured or discovered on a previous index), looking for links which it will follow breadth first, until there are no further pages to index or until the indexing page limit for the domain or the timeout is reached.
If you want to exclude certain pages from indexing, you would normally do this via robots.txt. If you have a Full listing you can also configure your listing to exclude content based on:
- path: i.e. URLs containing a certain string.
- type: i.e. values from the page_type field described above.
This might be useful to, for example, filter out micro blog entries which have a particular path or type.
How frequently are sites reindexed?
See Add Site for the latest information on indexing frequency. If you have a Full listing you can logon to Manage Site to see when your next reindex is due, and can of course trigger a reindex on demand.
How does the relevancy tuning work?
The following fields are used to determine how results are ranked: title, description, author, tags, url, content, indexed_inlink_domains_count, contains_adverts and owner_verified. There is further discussion of the relevancy tuning on some Blog posts, and of course the Source code is available for complete transparency.
Using the search
What is the query syntax?
Individual words: e.g. antarctica. If there are two words, e.g. antarctica book, it will search for them both but not as a phrase, and if there are three or more words, e.g. book about antarctica, it will search for a minimum of two of the words, e.g. in this example that could include pages with "book" and "about" but not "antarctica".
Phrase search: enclose phrase in double quotes to search for the exact phrase, e.g. "book about antarctica"
Boolean search: use AND, &&, NOT, !, OR, ||, + and -, with ( and ) to group queries, e.g. for pages which contain the keywords antarctica and book use antarctica AND book, or pages with antarctica and book but not movie use antarctica AND book !movie
Wildcard search: * for multiple characters, and ? for single characters, e.g. *arctic*
Filters: name:value, e.g. for all the pages on the michael-lewis.com domain which contain the word antarctica use domain:michael-lewis.com AND antarctica, or for all the article type pages on the michael-lewis.com domain use domain:michael-lewis.com AND page_type:article. See below for full list of field names.
Other searches: e.g. fuzzy searches, proximity searches, range searches, boost, etc. see The Standard Query Parser and The Extended DisMax Query Parser.
What fields are available?
| Name | Notes |
|---|---|
| id | URL of the web page, before following any redirects. Will be unique. |
| url | URL of the web page, after following any redirects. Will be the same as id if there are no redirects, and might not be unique if there are redirects. |
| domain | The domain to which the page belongs. |
| is_home | Boolean value, i.e. true or false. If true, indicates that the page is the home page for the domain. |
| title | Extracted from the title tag. |
| author | Extracted from meta name="author". |
| description | Extracted from meta name="description" or meta property="og:description". |
| tags | Multivalued. Extracted from meta name="keywords" or meta property="article:tag". |
| content | Text extracted from the main tag, or article tag, or body tag, with text from any nav, header and/or footer tags removed. |
| page_type | Extracted from meta property="og:type" or article data-post-type=. |
| page_last_modified | Extracted from the Last-Modified HTTP header. |
| published_date | Extracted from meta property="article:published_time" or meta name="dc.date.issued" or meta itemprop="datePublished". |
| date_domain_added | Date and time the domain was first added to the system for indexing. Only present on pages where is_home=true. |
| owner_verified | Boolean value, i.e. true or false. If true, indicates that the page is from a site which has been verified by the owner. |
| contains_adverts | Boolean value, i.e. true or false. If true, indicates that adverts have been detected on the page. |
| language | Extracted from html lang=. |
| language_primary | Language family, derived from the language attribute, e.g. if language=en-GB then language_primary=en. |
| indexed_inlinks | Multivalued. Pages which link to this page (from other domains within the search index, i.e. not from this domain or domains which aren't indexed). |
| indexed_outlinks | Multivalued. Pages to which this page links (to other domains within the search index, i.e. not to this domain or domains which aren't indexed). |
| indexed_inlink_domains | Multivalued. Unique domains in indexed_inlinks. |
Support
How do I raise a support query?
Use the Contact to reach out. If you think you have found a bug, you could also raise via https://github.com/searchmysite/searchmysite.net/issues.
How long will it take to get support?
As per About searchmysite.net, there is not a large team behind the service, so it may take a day or two to get a response. Note however that the service has been running reliably since July 2020.
Last updated: 22 September 2026.