Uploading Knowledge
Add documents and web content to a knowledge base so agents can retrieve from it at runtime. Xagent parses, chunks, embeds, and indexes everything you upload.
Upload Methods
- File upload — add one or many files at once, or drag and drop them onto the upload area.
- Website import — crawl a site starting from a URL and import the pages it finds.
- Cloud import — pull files directly from a connected cloud source (e.g. Google Drive) via
POST /api/kb/ingest-cloud. Native Google Slides presentations are exported to.pptxautomatically; Google Drive shortcuts are not supported. Cloud imports are capped at 5 files per request — split a larger batch into multiple calls.
Supported File Types
| Group | Formats |
|---|---|
| Documents | .pdf, .doc, .docx, .pptx, .txt, .md, .html |
| Data | .xlsx, .xls, .csv, .json |
| Code | .py, .js, and other common source formats |
File size
Maximum 100 MB per file. Split larger files before uploading.
Processing Options
When you upload, you control how documents are turned into searchable chunks:
| Option | Choices / default |
|---|---|
| Parse method | default (recommended), pypdf, pdfplumber, unstructured, pymupdf, deepdoc |
| Chunk strategy | recursive (default), fixed_size, markdown |
| Chunk size | characters per chunk (default 1000) |
| Chunk overlap | overlapping characters between chunks (default 200) |
Uploaded documents move through pending → processing → completed (or failed). Each is parsed, chunked, embedded, and indexed automatically.
Website Import
Point the crawler at a start URL and bound the crawl so it stays focused and respectful:
- Max pages (default 100) and crawl depth (default 3) cap how much is fetched.
- Concurrent requests (default 3, max 10) and request interval (default 1s) throttle load on the target.
- Timeout (default 30s) bounds how long each page fetch may take — raise it for slow sites.
- URL / exclude patterns, same-domain only, and CSS content / remove selectors filter what is kept.
- Follow robots.txt keeps crawling within the site's stated policy.
Tune chunking to your content
Technical docs benefit from larger chunks with more overlap (~1500–2000 / 300–500). FAQs work better with smaller chunks (~500–800 / 100–200). For code, use the markdown strategy to preserve structure.
Troubleshooting
| Problem | What to do |
|---|---|
| The document is refused as an unsupported type | Check the extension against Supported File Types above. The type is checked on upload, so renaming the file will not get it through — convert it to a supported format. |
| The document is refused even though its type is supported | The parse method you chose may not work for that file. Xagent checks the file and the requested parser together and rejects the pair rather than producing an unusable result. Leave the parse method on its default and let Xagent choose — see Processing Options. |
| The upload is refused with a message about upgrading | Your team has reached the total knowledge base storage on its plan. Deleting files frees capacity immediately — see Storage Limits. |
| The file uploaded, but agents do not use it | Uploading only adds the document to the knowledge base. An agent uses it only once that knowledge base is attached to the agent in its configuration — see Building Agents. |
Next Steps
- Knowledge Retrieval → — search and use what you uploaded
- Knowledge Bases Overview →
- Embedding Models →