Chunk text with overlap
Before indexing a document for RAG you split it into chunks. A small overlap avoids cutting an idea right in half.
Write chunkWords(text, size, overlap). Split text into words on any whitespace (spaces, newlines, tabs) and return an array of strings, each with at most size words joined by one space. Each chunk starts size - overlap words after the previous one, so it shares overlap words with it.
Do not produce empty chunks or a final chunk that only repeats words from the previous one. The last chunk may be shorter. An empty text returns []. Throw an error when size is not a positive integer or overlap is not between 0 and size - 1.
DOC holds a 12-word example text.
Challenges 0/4
- Splits DOC into chunks of 4 with an overlap of 1
- Leaves no final chunk that only repeats the overlap
- With no overlap, the last chunk may be shorter
- Empty text gives [] and an invalid overlap throws
function chunkWords(text, size, overlap) {
// split on whitespace, then take `size` words every `size - overlap` words
const words = text.split(' ');
const chunks = [];
for (let i = 0; i < words.length; i += size) {
chunks.push(words.slice(i, i + size).join(' '));
}
return chunks;
}
console.log(chunkWords(DOC, 4, 1));Go deeper: the related guide →
This in production, with your data? Let's talk for 15 minutes →