Skip to main content

The file connector supports CSV, Excel and JSON files.

Tabset

Read​

Recipe​

read:
- file:
name: file.csv

# Optional
nrows: 10 # Limit the number of rows
columns:
- column1
- column2

Function​

from wrangles.connectors import file
df = file.read('file.csv')

Parameters​

ParameterRequiredData TypeNotes
name✓strThe file name (and path, if required) to read.
columnslistA list with a subset of the columns to import.
not_columnslistSubset of columns to be left out of the read.
decimalstrUsed for CSV files. Character to recognize as the decimal point (e.g. ',' for European data).
encodingstrUsed for CSV files. Set the encoding used for the file. Default utf-8.
file_objectBytesIOFunction Only. Pass in a file object from memory instead of reading from the file system. If this is provided a name is still required to indicate the file type, but won't be read.
headerintSet the header row number.
nrowsintLimit the number of rows.
orient(str) - split / records / index / columns / valuesUsed for JSON files. Specifies the input arrangement. See pandas docs for details
sepstrUsed for CSV files. Set the separation character. Default , (comma).
sheet_namestrUsed for Excel files. Specify the sheet to read.
thousandsstrUsed for CSV files. Character to recognize as the thousands separator.
order_bystrUses SQL syntax to sort the input.
ifstrA condition that will determine whether the action runs or not as a whole.

Write​

Recipe​

write:
- file:
name: file.xlsx

# Optional
columns:
- column1
- column2

Function​

from wrangles.connectors import file
file.write(df, 'file.xlsx')

Parameters​

ParameterRequiredData TypeNotes
name✓strThe file name (and path, if required) to write.
columnslistSubset of the columns to be written. If not provided, all columns will be output.
not_columnslistSubset of columns to be left out.
decimalstrUsed for CSV files. Character to recognize as the decimal point (e.g. ',' for European data).
encodingstrUsed for CSV files. Set the encoding used for the file. Default utf-8.
file_objectBytesIOFunction Only. Pass in a file object from memory instead of reading from the file system. If this is provided a name is still required to indicate the file type, but won't be read.
headerintSet the header row number.
indexbooleanInclude a column with the row index in the output. Default false.
modestrUsed for CSV files. Set whether to append to (a) or overwrite (w) the file if it already exists. Default w - overwrite
nrowsintLimit the number of rows.
orient(str) - split / records / index / columns / valuesUsed for JSON files. Specifies the input arrangement. See pandas docs for details
sepstrUsed for CSV files. Set the separation character. Default , (comma).
sheet_namestrUsed for Excel files. Specify the sheet name.
order_bystrUses SQL syntax to sort the output.
ifstrA condition that will determine whether the action runs or not as a whole.