Website Knowledge Sync
Integrate your team's knowledge directly from your website into Brainfish. By syncing your website content, you can keep your AI agent's knowledge up-to-date automatically, ensuring it always provides the most accurate and relevant answers based on your website documentation, blog posts, and other content.
This guide will walk you through the process of connecting and syncing your website as a knowledge source.
Step 1: Navigate to External Data and add a Website source
- From your Brainfish Dashboard, open the main navigation.
- Select External Data.
- Click + Source.
- Select Website (it is the first option in the list).
Step 2: Configure Your Website Connection
You will now see the "Configure Website" form. You'll need to provide a name for this source and the URLs you want to sync.
Source Name: Enter a descriptive name for your connection, such as "Company Website" or "Product Documentation".
URLs to Import: Add the URLs you want to sync from your website. You can add multiple URLs by clicking the + Add URL button.
How to Add URLs
- Single Page: Enter the full URL of a specific page (e.g.,
https://yourcompany.com/about) - Multiple Pages: Add each page URL separately
- Crawl Options: For each URL, you can choose whether to:
- Crawl the website: Automatically discover and sync linked pages from this URL
- Single page only: Only sync the specific URL you entered
URL Format Requirements
- Protocol: URLs should start with
https://orhttp:// - Valid Domain: Ensure the domain is accessible and not behind authentication
- Public Content: Only public pages can be crawled (no login-required pages)
Advanced Configurations (Super Admin only):
- Extract Images: Toggle to include images in the synced content
- Extract Iframe Content: Toggle to include content from iframes
Click the + Add Source button to finalize the connection.
Step 3: Sync and Verify Your Content
After adding the source, Brainfish will begin importing content from your website.
- On the External Data page, you will see your new website source with a "Syncing" status. This process may take several minutes, depending on the number of pages and complexity of your website.
- Once the sync is complete, the status will update to show the number of pages imported.
- Click on the newly created website catalog to view the list of imported pages. You can click the eye icon (👁️) next to any page to preview its content and ensure it was imported correctly.
Step 4: Enable the Source for Your Agent
To start using this knowledge, you must enable the source and ensure your agent has access to it.
- On the External Data page, find your new website source card. Click the three-dots menu (...) and ensure the source is Enabled.
- Navigate to the Agents page from the left-hand menu.
- Select the agent you want to train with this new knowledge.
- In the agent configuration panel, go to Knowledge Settings.
- Ensure that All sources is selected, or specifically add your new website source from the dropdown menu.
Step 5: Test Your Integration
Finally, test that your agent can find and use information from the new source.
- Open the Agent Playground or your website where the agent is deployed.
- Ask a question related to the content on your synced website pages.
- The agent should provide a detailed answer based on the information from your website content. You can also expand the Sources to see the specific webpage that was used to generate the answer.
You have now successfully synced your website with Brainfish! Your agent will now be able to answer questions using the knowledge stored on your website.
Troubleshooting
Common Issues
"URL is not valid" error
- Verify the URL format is correct (starts with http:// or https://)
- Check that the URL is accessible from the internet
- Ensure the domain is not misspelled
"Page not found" error
- Confirm the URL exists and is publicly accessible
- Check if the page requires authentication
- Verify the page is not behind a paywall or login
"Crawl limit exceeded" error
- Trial accounts are limited to 50 pages per website
- Consider upgrading your plan for higher limits
- Focus on the most important pages for your use case
"Content extraction failed" error
- The page might be using complex JavaScript that requires rendering
- Check if the page has anti-bot protection
- Ensure the page contains readable text content
"Request timeout" error
- Large websites may take longer to crawl
- Check your internet connection
- Try again during off-peak hours
Best Practices
- Start Small: Begin with a few key pages before crawling entire websites
- Focus on Quality: Prioritize pages with valuable, relevant content
- Regular Updates: Set up daily sync to keep content current
- Monitor Performance: Check sync status and imported content regularly
Content Optimization Tips
- Clear Structure: Ensure your website has clear headings and well-structured content
- Relevant Content: Focus on pages that contain useful information for your users
- Avoid Duplicates: Don't add multiple URLs that point to the same content
- Update Regularly: Keep your website content fresh and up-to-date
Support
If you encounter any issues during setup, please contact our support team with:
- The specific URLs you're trying to sync
- The error message you're seeing
