Ulabase

AI

Auto Chunking

Rules on a file bucket that split the files uploaded to it into text chunks, in a collection you choose; that collection embeds them, and you search it like any other.

Under AI → Auto Chunking every file bucket of the service is listed, with its chunking rules. A rule says which files it takes, the collection their chunks go to, and how big a chunk is. From then on every file uploaded to the bucket, by your app or by the console, is split by the first rule it matches. A bucket without rules is never touched.

The chunks get no vector from the rule. Their collection embeds them with its own rule, set under Auto Embeddings, like any other collection: one place for embedding, the collection where the vectors live and where you search.

The rule

Open a bucket and Add chunking rule:

  • Name: written on every chunk as rule.

  • Condition: which files the rule takes, below. Empty: every file.

  • Chunks into: the collection of the chunks, <bucket>_chunks by default, created empty on the first file if missing.

  • Advanced: chunk size in characters, 1000 by default; the overlap taken from the chunk before, 200; the splitter, auto to cut source code at function and class boundaries and the rest as prose, text for prose always, code for boundaries whenever the language is known.

Prose is cut at paragraphs first, then at lines, sentences and words when a paragraph is longer than a chunk, and the pieces are put back together as full as the chunk size allows; the overlap is whole sentences or words of the chunk before. A word is never cut. A blank page makes no chunk, and a chunk with fewer than 20 letters and digits, a page number or a running header, is joined to its neighbour.

Save writes the whole list on the bucket, as chunking in its metadata. The rules are tried in the order of the list, and the first one a file matches is applied, as in an ACL: the arrows move a rule up or down.

The condition

A JSON object with three keys, each optional; when there are more, all must match.

Key Matches

contentType

the type detected from the file’s content, a string or a list; a trailing matches any suffix: "application/vnd.openxmlformats-officedocument.".

extension

the extension of the file’s name, a string or a list, with or without the dot.

metadata

a MongoDB filter on the file’s metadata, the document your app sends with the upload. Every query operator works but those that run code.

The quick buttons under the field write the common ones: PDF, Office, Text, Markdown, HTML, Source code, By folder, By metadata.

By folder. A bucket has no folders: a folder is part of the path, which your app sends in the file’s metadata.path, next to the name in metadata.filename. Two rules on a bucket docs send the contracts and the design documents to two collections:

{ "chunking": [
    { "name": "legal",
      "filter": { "metadata": { "path": { "$regex": "^legals/" } } },
      "target-collection": "legal_chunks" },
    { "name": "tech",
      "filter": { "extension": ".pdf", "metadata": { "path": { "$regex": "^design/" } } },
      "target-collection": "tech_chunks" } ] }

legals/contratto.pdf goes to legal_chunks, design/project.pdf to tech_chunks. The condition sees only what is in metadata: upload with the path there.

curl -X POST https://<service>/docs.files \
  -H "Authorization: Bearer <token>" \
  -F file=@contratto.pdf \
  -F 'metadata={"filename": "contratto.pdf", "path": "legals/contratto.pdf"}'

By metadata. A field of your own does the same without a folder: {"metadata": {"area": "legal"}}, uploading with metadata={"filename": "contratto.pdf", "area": "legal"}.

Under the rule

  • Chunks: how many the rule wrote.

  • Embedding: the rule of the chunks collection, if it has one, with a link to Auto Embeddings. Without one the chunks have no vector: Add embedding opens Auto Embeddings on that collection with text into vector filled in and, for a contextual model, the chunks of a file grouped by fileId, so each vector knows the rest of its document.

Each chunk carries text, chunkIndex, fileId, source, filename, contentType, the file’s metadata and rule. Replacing or deleting a file replaces or deletes its chunks.

In the background

The upload answers as soon as the file is stored, and the file is split and embedded afterwards: a large PDF takes seconds. The file’s own document says where it is, under chunking: pending right after the upload, then done with the rule applied and the number of chunks, skipped when no rule matched, or failed with the reason in warnings.

GET /docs.files?filter={"_id":{"$oid":"<id>"}}

A service chunks one file at a time on a free plan, two on a shared one; the others wait their turn.

Under the rules of a bucket, Files shows its queue: how many files are waiting, running, done, skipped and failed, and the files waiting or running, oldest first, with the ones that failed and why. While something is waiting or running the list refreshes on its own, and the row of the bucket says how many are in the queue.

Upload a file to try

Under the rules, choose a file and, if you want, a folder: the file is uploaded to the bucket with its name in metadata.filename and the folder and the name in metadata.path, as your app would. The page waits for the chunking to finish, then says which rule the file matched, shows its chunks and whether they have a vector, and repeats the warnings. The file stays in the bucket: Delete the file and its chunks removes both.

Upload a folder

Under a bucket, Upload a folder shows a Python script with the service’s URL and the bucket already in it. It uploads every file under a directory, subfolders included, each with its name in metadata.filename and its path from the directory in metadata.path. Put a token in place of <token> and run it with the directory:

python upload.py ./documents

Limits

The chunks of a file are written in one request, with the limits of any request. On a free service a file is up to 2 MB, and so are its chunks together: a file whose chunks exceed it is stored, its chunks are not, and the warning of the upload says so. Free is for trying; for a real archive, a shared or dedicated service. The chunks and their vectors count toward the storage of the plan.

What it costs

Splitting is free. Embedding the chunks is billed by your provider to your account, one embedding per chunk, or one call per file with a contextual model grouped by fileId.