ContextPack
A command-line tool that packs the files you choose into a fixed token budget, keeps the required ones whole, and writes a report explaining every section it kept or dropped.
- Context
- Personal developer tool.
- Stack
- Python 3.11 or newer, tiktoken 0.14.0 for GPT-2 token counting, tomllib for configuration, pytest for tests. Optional extra: PyTorch for a small character-level model, which the packing path never imports.
- Role
- Sole author
- Repository
- github.com/HeyItWorked/ContextPack (Public)

At a glance
What it is
A language model will only read so much at once. When I wanted to hand one a bug, I kept doing the same thing by hand: paste the traceback, paste the config, add as much of the log and the source as would fit, and guess at the rest.
ContextPack does that deliberately. A small TOML file names the task, a token budget, and two lists of files. Required files are kept whole. Optional files are split into sections and packed into whatever space is left. The result is one text file to paste, and a report saying what did not make it.
What is implemented
- A
buildcommand that reads a TOML file and writespacked-context.txtandreport.json. - Required files are kept whole; if they do not fit the budget, the build fails rather than truncating them.
- Optional files are split into fixed 50-line windows and tried in the order the files are listed.
- Every candidate is counted as a complete document, not as a sum of its pieces, so headings, separators and token merges at the boundaries are inside the budget rather than beside it.
- The report records the schema and tokenizer versions, the budget, the final count, a SHA-256 of the output, and a reason for every section kept or dropped.
- A separate
generatecommand runs a small character-level Transformer over an existing packed file. No checkpoint ships with the project, and the packing path never imports PyTorch.
Verification
Checked on 21 Sep 2026 on Python 3.12, with the package installed from the repository.
- The full suite is 57 tests and all of them pass. 26 run without PyTorch installed, which confirms the packing path really is independent of it.
- The worked example in the README reproduces exactly:
271 / 2000tokens, 4 sections in, 0 out. - Determinism holds. Two runs over the same inputs produced files with the same SHA-256, and that hash matched the
output_sha256the report had written for itself. - Dropping the same example to a 120-token budget kept the two required files and excluded both optional windows, each with its line range and the reason does not fit remaining budget.
These are my own measurements from running the tool, not figures quoted from its README.
Limitations
- Counting uses GPT-2 rules through tiktoken. A different model will divide the same text differently, so the budget is an approximation for anything that is not GPT-2.
- Optional sections are tried in the order you list them. There is no scoring or relevance ranking, and picking the right files is still your job.
- It does not search a repository or detect secrets. Whatever you point it at goes in.
- A failed build writes a report but leaves any earlier packed file in place.
- The character-level model is a learning exercise carried over from a separate project. It is not a summarizer or a debugger, and it needs a checkpoint you supply yourself.
- Version 0.1.0, not published to PyPI.