<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Csvkit on josiete.com</title><link>https://www.josiete.com/en/tags/csvkit/</link><description>Recent content in Csvkit on josiete.com</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 07 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.josiete.com/en/tags/csvkit/index.xml" rel="self" type="application/rss+xml"/><item><title>csvkit: a toolkit for working with CSV files in the terminal</title><link>https://www.josiete.com/en/posts/csvkit/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate><guid>https://www.josiete.com/en/posts/csvkit/</guid><description>&lt;p&gt;&lt;img src="https://www.josiete.com/images/csvkit.webp" alt="Illustration of a terminal connected to CSV documents and data tables through processing stages"&gt;&lt;/p&gt;&#10;&lt;p&gt;A CSV file arrives. Before loading it into a database, you want to see its columns, check for missing values, and find the records you need. Opening a spreadsheet works, but becomes inconvenient when you have to repeat the same operation across several files. Writing a program for every check is not always worthwhile either.&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://csvkit.readthedocs.io/en/latest/"&gt;csvkit&lt;/a&gt; brings together command line tools for converting and processing tabular data. It is written in Python and handles many of these tasks with short commands you can save and run again.&lt;/p&gt;&#10;&lt;h2 id="csv-has-more-structure-than-it-seems"&gt;CSV has more structure than it seems&lt;/h2&gt;&#10;&lt;p&gt;Splitting a line on commas might seem sufficient until a value such as &lt;code&gt;&amp;quot;Madrid, downtown&amp;quot;&lt;/code&gt; appears. Fields may also contain escaped quotes or line breaks. That is why using &lt;code&gt;cut&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;, or &lt;code&gt;sort&lt;/code&gt; directly on the text requires care: those tools do not interpret CSV structure on their own.&lt;/p&gt;&#10;&lt;p&gt;csvkit works with rows and columns. You can select a column by name, filter a particular field, and connect operations with &lt;code&gt;|&lt;/code&gt;. Transformation commands still output CSV; &lt;code&gt;csvlook&lt;/code&gt;, however, produces a table for reading on screen. The &lt;a href="https://csvkit.readthedocs.io/en/latest/tutorial/1_getting_started.html"&gt;official tutorial&lt;/a&gt; explains this use of standard input and output.&lt;/p&gt;&#10;&lt;h2 id="installation"&gt;Installation&lt;/h2&gt;&#10;&lt;p&gt;The documentation recommends a virtual environment. On Linux or macOS:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python3 -m venv .venv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;source .venv/bin/activate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python -m pip install csvkit&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvcut --version&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;In PowerShell, activate it with &lt;code&gt;.venv\Scripts\Activate.ps1&lt;/code&gt;. Installation provides several executables, such as &lt;code&gt;csvcut&lt;/code&gt; and &lt;code&gt;csvstat&lt;/code&gt;; there is no single command named &lt;code&gt;csvkit&lt;/code&gt; to invoke.&lt;/p&gt;&#10;&lt;h2 id="a-reproducible-example"&gt;A reproducible example&lt;/h2&gt;&#10;&lt;p&gt;Save this UTF-8 content as &lt;code&gt;sales.csv&lt;/code&gt;:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-csv" data-lang="csv"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;id&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;customer&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;category&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;units&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;price&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;status&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;1&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;&amp;#34;North Bookshop&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;books&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;18.50&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;paid&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;&amp;#34;Cafe, Central&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;home&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;1&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;32.00&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;pending&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;3&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;&amp;#34;North Bookshop&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;books&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;3&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;12.00&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;paid&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;4&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;&amp;#34;South Studio&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;home&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;25.00&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;paid&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The comma in &lt;code&gt;&amp;quot;Cafe, Central&amp;quot;&lt;/code&gt; belongs to the customer name; it does not separate two columns.&lt;/p&gt;&#10;&lt;h3 id="inspect-before-transforming"&gt;Inspect before transforming&lt;/h3&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvcut -n sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvlook sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvstat sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvstat --count sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;csvcut -n&lt;/code&gt; lists columns and &lt;code&gt;csvlook&lt;/code&gt; makes the data easier to read. &lt;a href="https://csvkit.readthedocs.io/en/latest/scripts/csvstat.html"&gt;csvstat&lt;/a&gt; summarizes inferred types, null values, and statistics appropriate to each column type. The final command returns &lt;code&gt;4&lt;/code&gt;: it counts records, not physical file lines. This matters when fields contain line breaks.&lt;/p&gt;&#10;&lt;h3 id="select-columns-and-filter-rows"&gt;Select columns and filter rows&lt;/h3&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvgrep -c status -r &lt;span style="color:#e6db74"&gt;&amp;#39;^paid$&amp;#39;&lt;/span&gt; sales.csv &lt;span style="color:#ae81ff"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; | csvcut -c customer,category,units,price &lt;span style="color:#ae81ff"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &amp;gt; paid_sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvsort -c price -r paid_sales.csv | csvlook&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;a href="https://csvkit.readthedocs.io/en/latest/scripts/csvgrep.html"&gt;csvgrep&lt;/a&gt; applies the regular expression to the &lt;code&gt;status&lt;/code&gt; column. The &lt;code&gt;^&lt;/code&gt; and &lt;code&gt;$&lt;/code&gt; anchors match the entire field. We then select four columns and save the result. &lt;code&gt;csvsort&lt;/code&gt; sorts by price from highest to lowest; here prices are interpreted as numbers.&lt;/p&gt;&#10;&lt;p&gt;A pipe connects one command&amp;rsquo;s output to the next command&amp;rsquo;s input. &lt;code&gt;&amp;gt;&lt;/code&gt; writes that output to a file, overwriting its contents if it already exists. Use a different name from the input file to avoid truncating it before it is read.&lt;/p&gt;&#10;&lt;h3 id="aggregate-with-sql"&gt;Aggregate with SQL&lt;/h3&gt;&#10;&lt;p&gt;To calculate paid sales revenue by category:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvsql --query &lt;span style="color:#e6db74"&gt;&amp;#34;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; SELECT category,&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; SUM(units * price) AS revenue&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; FROM sales&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; WHERE status = &amp;#39;paid&amp;#39;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; GROUP BY category&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; ORDER BY revenue DESC&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt; sales.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The result totals &lt;code&gt;73&lt;/code&gt; for books and &lt;code&gt;50&lt;/code&gt; for home; decimal formatting may vary. By default, the table name comes from the filename without its extension: &lt;code&gt;sales.csv&lt;/code&gt; becomes &lt;code&gt;sales&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://csvkit.readthedocs.io/en/latest/scripts/csvsql.html"&gt;csvsql&lt;/a&gt; loads the data into an in-memory SQLite database to run this query. It can also generate SQL statements or import data into a database. For monetary calculations requiring exactness, define a precision strategy: SQL execution can introduce floating point effects.&lt;/p&gt;&#10;&lt;h2 id="join-files-and-change-formats"&gt;Join files and change formats&lt;/h2&gt;&#10;&lt;p&gt;Suppose &lt;code&gt;customers.csv&lt;/code&gt; contains:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-csv" data-lang="csv"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;customer&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;city&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;North Bookshop&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Bilbao&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Cafe, Central&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Madrid&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;South Studio&amp;#34;&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Seville&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We can add the city to each sale:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvjoin --left -c customer sales.csv customers.csv &amp;gt; sales_with_city.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;a href="https://csvkit.readthedocs.io/en/latest/scripts/csvjoin.html"&gt;csvjoin&lt;/a&gt; combines tables using a key. &lt;code&gt;--left&lt;/code&gt; keeps every row in the first file, even without a match in the second. If a key appears several times in both files, the join can multiply rows: check that this relationship is the one you need.&lt;/p&gt;&#10;&lt;p&gt;For other inputs and outputs:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;in2csv sales.xlsx &amp;gt; sales_from_excel.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvjson sales.csv &amp;gt; sales.json&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvformat -D &lt;span style="color:#e6db74"&gt;&amp;#39;;&amp;#39;&lt;/span&gt; sales.csv &amp;gt; sales_semicolon.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;in2csv&lt;/code&gt; converts an Excel worksheet into tabular data; it does not preserve its presentation. &lt;code&gt;csvjson&lt;/code&gt; produces JSON and &lt;a href="https://csvkit.readthedocs.io/en/latest/scripts/csvformat.html"&gt;csvformat&lt;/a&gt; changes the output delimiter. Notice the difference: &lt;code&gt;-d&lt;/code&gt; specifies the input separator, while &lt;code&gt;-D&lt;/code&gt; in &lt;code&gt;csvformat&lt;/code&gt; specifies the output separator.&lt;/p&gt;&#10;&lt;h2 id="delimiters-encodings-and-types"&gt;Delimiters, encodings, and types&lt;/h2&gt;&#10;&lt;p&gt;For a semicolon-delimited file encoded in Windows-1252:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;csvlook -d &lt;span style="color:#e6db74"&gt;&amp;#39;;&amp;#39;&lt;/span&gt; -e cp1252 -y &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt; -I export.csv&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;-y 0&lt;/code&gt; disables automatic dialect detection. &lt;code&gt;-I&lt;/code&gt; disables type inference in commands that support it, such as &lt;code&gt;csvlook&lt;/code&gt;, &lt;code&gt;csvstat&lt;/code&gt;, and &lt;code&gt;csvsql&lt;/code&gt;. This is useful for keeping identifiers such as &lt;code&gt;00123&lt;/code&gt; as text. The &lt;a href="https://csvkit.readthedocs.io/en/latest/common_arguments.html"&gt;common arguments&lt;/a&gt; document encoding and format options; also check the individual command&amp;rsquo;s help.&lt;/p&gt;&#10;&lt;p&gt;Think about the next operation before disabling inference: a price treated as text may sort differently from a number. Successfully reading a file does not establish that its data is valid; business rules still need their own checks.&lt;/p&gt;&#10;&lt;h2 id="where-it-fits"&gt;Where it fits&lt;/h2&gt;&#10;&lt;p&gt;csvkit is convenient for exploring an export, preparing a load, or automating small, repeatable transformations. You can save the commands in a script and review exactly which filters you applied.&lt;/p&gt;&#10;&lt;p&gt;As volume or complexity grows, consider an engine such as &lt;a href="https://www.josiete.com/en/posts/duckdb/"&gt;DuckDB&lt;/a&gt; or a library such as &lt;a href="https://www.josiete.com/en/posts/polars-python/"&gt;Polars&lt;/a&gt;. Some csvkit operations need to hold data in memory, and each pipeline command may parse it again. Its main appeal is convenience: turning an unfamiliar CSV into a clear transformation with a handful of instructions.&lt;/p&gt;&#10;</description></item></channel></rss>