Capstone: from a messy CSV to a summary you can trust
Puts the track together: parse CSV with quotes, validate every field with its line number and aggregate per model.
Combines the 6 exercises of the Data and text track. It unlocks when you finish them, but you can try it now.
RUNS is a speed-test log as people really keep it: quoted notes, a blank line, typos and \r\n line ends. Write summarize(csv) that returns { valid, errors, byModel }.
1. Split into lines (\n or \r\n). Line 1 is the header. Skip blank lines, but count them for numbering. Fields may be quoted with commas inside and "" as a quote; there will be no line breaks inside a field.
2. Validate each row: model not empty, date a real YYYY-MM-DD date and tokens_per_sec a number above 0. For every field that fails, add { line, field } to errors, in column order. A row with any error does not count.
3. valid is the number of good rows. byModel is a list of { model, runs, avgTps } sorted by model, with the average rounded to one decimal. Use that key order.
Challenges 0/5
- Counts 5 valid rows (quotes do not break the row)
- Reports each error with its line and field
- Aggregates per model: runs and the average to one decimal
- A row with three problems gives three errors, in column order
- Copes with empty text and a header alone
function summarize(csv) {
// 1. split into lines (\n or \r\n); line 1 is the header; skip blank lines but keep counting them
// 2. split each line into fields: a comma inside "quotes" does not split, and "" is a quote
// 3. check model (not empty), date (a real YYYY-MM-DD) and tokens_per_sec (a number above 0)
// 4. average tokens_per_sec per model over the valid rows, one decimal, sorted by model
return { valid: 0, errors: [], byModel: [] };
}
console.log(summarize(RUNS));Go deeper: the RFC 4180 reference →
This in production, with your data? Let's talk for 15 minutes →