Agents

Crawling websites into a knowledge base

Add a website to a knowledge base, control how far the crawl goes, and keep it fresh on a schedule

Overview

Add URL crawls a website and adds it to a knowledge base. MagOneAI walks same-domain links from the URL you give it, collects each page as text, and stores the whole crawl as one document in the knowledge base.

One Add URL action creates one row in the document list, not one row per page. Page boundaries are preserved inside the document, so a citation still points at the page it came from.

You can also give that document a schedule, so the site is re-crawled on a cadence and the knowledge base stays current without anyone remembering to refresh it.

Adding a URL

The Add from URL panel in a MagOneAI knowledge base, showing the seed URL field, the max pages slider set to 150, the max depth selector set to 2, and the re-crawl on a schedule checkbox
Setting up a crawl in the Add from URL panel.

Open the knowledge base and click Add URL

The Add from URL panel opens. Everything the crawl needs is on this one panel, including the schedule, so a recurring crawl is set up in a single pass.

Enter the seed URL

Seed URL is the page the crawl starts from. MagOneAI renders it in a headless browser, so a site that builds its content with script is still read correctly.

Public URLs only. See What is crawled, and what is refused.

Set the page budget

Max pages caps the total number of pages the crawl fetches. The slider snaps to 10, 25, 50, 100, 150, 250 and 500, and starts at 150.

The panel shows an estimated run time under the slider, at roughly three seconds per page. The low stops are there so you can test a crawl on a small site before you commit to a large cap.

Set the depth

Max depth is how many links away from the seed the crawl may travel. The panel offers five choices, and starts at 2.

DepthCrawls
0The seed page only
1The seed, and every same-domain page it links to
2The above, plus pages those link to
3One level further again
UnlimitedEvery same-domain link, until the page budget runs out

The default of 2 reaches a site's main content without spending the page budget on deep or archived pages.

Add a schedule, if you want one

Tick Re-crawl on a schedule to keep the document current. See Scheduling a recurring crawl.

Click Start Crawl

The crawl runs in the background and the document appears as processing. Crawling never happens inside the request, so a large site does not hold a request open.

The API accepts a wider range than the panel offers: up to 500 pages and a depth of up to 10. A value above either ceiling is clamped down rather than refused.

Depth and the page budget work together

The crawl walks breadth first, so it finishes everything at one depth before going deeper.

That matters because the two limits interact. With a budget of 150 pages and unlimited depth on a large site, the budget is spent near the seed rather than following one deep chain of low-value pages. Setting a depth as well is how you say "the documentation section, not the whole site".

Start with a low depth and a small page budget, look at what arrived, then widen. A depth of 1 on a documentation index page is often all you need, and it finishes in a fraction of the time.

What is crawled, and what is refused

RuleBehaviour
Links followedSame domain as the seed only
SchemeHTTPS only by default
AddressesOnly publicly routable ones. Private and internal addresses, and cloud metadata hostnames, are refused.
RedirectsThe final URL is re-checked after every redirect
Error pagesA page returning 400 or above is skipped
Bot protectionDetected and skipped rather than stored as content
Page sizeCapped, 10 MB by default
PolitenessA short delay between pages, 0.5 seconds by default

The crawler identifies itself and respects the page budget, but a scheduled crawl sends unattended traffic to someone else's site. Point it at sites you own or are permitted to crawl, and keep the cadence modest.

Crawling is refused for private and internal addresses. This is deliberately stricter than other parts of MagOneAI, because a crawler takes a URL from a user rather than from an administrator.

Refreshing a crawled document

Each crawled document has a refresh control that re-crawls the same URL end to end, reusing the page and depth settings it was created with.

The refresh is content-aware:

  • If the new content is identical, the existing vectors are left in place. Nothing is re-embedded.
  • If it changed, the old vectors are removed and the new content is indexed.

Re-adding a URL whose crawl failed reuses the existing row rather than creating a second one, so a retry does not leave two entries for the same site.

Scheduling a recurring crawl

Tick Re-crawl on a schedule on the Add from URL panel, and the document is re-crawled on the cadence you pick.

Pick a frequency

Three presets cover most cases:

PresetRuns
HourlyAt the top of every hour
DailyOnce a day at 3am
WeeklyEvery Monday at 3am

Choose Custom for anything else.

Write a cron expression, for a custom cadence

Standard five-field cron: minute, hour, day, month, weekday. For example 0 3 * * 1 for 03:00 every Monday.

Set the timezone

The cadence is read in the timezone you pick, so "3am" means local business hours rather than UTC.

Save

The document shows its schedule and its next run time. The next run time is calculated by the server, because the cron is evaluated in the timezone you chose.

The same panel reopens as Edit crawl settings on a document that already has a schedule, so the page cap, the depth and the cadence are all changed in one place.

The minimum interval is one hour

A cadence with any gap shorter than an hour is refused, not quietly slowed down.

Refusing rather than clamping is deliberate. Someone who asked for every five minutes and silently got hourly would have no way to tell. The error states the floor.

The floor exists because a crawl is much heavier than an ordinary workflow: fetching the pages is one long job, and each finished crawl then triggers a full extract, chunk, embed and index pass on the same worker pool.

Note that the floor applies to the shortest gap in the pattern, not the average. An expression like 0,5 * * * * runs twice an hour but with a five-minute gap, so it is refused.

Limits

LimitDefault
Minimum gap between runs1 hour
Documents with a schedule, per project25
Scheduled crawls started per sweep10

Scheduled crawls are checked every 15 minutes, and the ones due longest are started first. If more are due than one sweep will start, the rest are picked up by the next sweep rather than dropped.

Pausing and removing a schedule

Tick Pause for now on the schedule to stop the runs without losing the configuration, and untick it later to resume. Clearing Re-crawl on a schedule leaves the document and its content in place; only the cadence goes.

Troubleshooting

Next steps

MagOneAI© 2026 Magure, Inc.

On this page