Calculate effective chunk stride and overlap ratio for document splitting in RAG.
Chunk overlap ensures that information spanning chunk boundaries is captured in at least one chunk. The stride (chunk_size - overlap) determines how far the window advances. More overlap means more chunks but better retrieval recall at chunk boundaries.
Chunk Count
chunks = ceil((doc_length - overlap) / (chunk_size - overlap))
10-20% overlap is a good default. Higher overlap (25-50%) improves recall for sentence-level retrieval but increases storage and embedding costs proportionally.
Yes — more chunks means more embedding API calls. With 20% overlap you'll have ~25% more chunks than with no overlap, increasing embedding costs accordingly.