SI Data Ops

One task · £195 · untested offer price

Make a supplier CSV import keep accented and special characters intact

A named supplier CSV with accents, symbols and a leading byte-order mark imports with every agreed test value unchanged and its first column recognised; before and after results attached.

Request this outcome

Who this is for

Operations or e-commerce manager at a distributor or online shop whose product import shows garbled names, or cannot find its first column.

Product names show character pairs such as é or the replacement sign � after an import, or the importer says the SKU column is missing although the header looks right.

The result

A named test file with agreed non-ASCII values and a leading byte-order mark imports through your existing importer with each value identical to the source and the first column recognised. You receive the repair as a reviewable change with before and after results.

What is included

  • One existing importer path that reads one named supplier CSV layout
  • Reproduce the fault on a synthetic file built from your redacted sample, recording the exact bytes of the header and of up to ten affected values
  • Make the importer decode the file with an explicit, agreed encoding named in the code rather than a language or system default, handle a leading byte-order mark, and reject bytes that are not valid in that encoding instead of storing substitutes
  • Add a regression test that uses the synthetic file

You receive

  • A change to the importer with its regression test
  • A before and after table of expected and stored values for the agreed test rows, shown as code points
  • A short note on which encodings the importer now accepts and what it does with undecodable bytes
  • Steps to revert the change

What is not included

  • Repairing text already stored garbled in your database; that is a separate reconciliation job
  • Recovering the original characters from a file that an earlier spreadsheet save already damaged
  • Guessing the encoding of arbitrary files from unknown suppliers
  • Changing database character sets or collations
  • Faults caused by delimiters, quotes or line breaks inside fields, which are a different job

What we need from you first

  • Three to ten invented or redacted rows that show the fault, as plain text, including the header line
  • The importer's name and language, and which program the supplier says exports the file
  • What you see: the wrong characters or the missing-column message. No real price lists, customer data, credentials or code in the first enquiry

Never send passwords, keys, customer records or confidential code in the first enquiry. Secure handover is agreed after scoping.

How we check it is done

  • A synthetic file that starts with a byte-order mark and contains the agreed non-ASCII values imports with every stored value equal to the expected value and the first header name equal to the agreed column name.
  • The same file saved without a byte-order mark imports to identical values.
  • A file containing bytes that are not valid in the agreed encoding is rejected with a message that names the line, and no rows from it are stored.
  • The importer's existing tests still pass with the change applied.

You inspect the before and after table and the test results, and sign off in writing. Payment follows sign-off; applying the change to your live importer stays with your maintainer.

When we would stop or decline

  • The only reproduction needs production data, a live supplier download or credentials
  • The original characters are already lost in the file you hold, so no decoding can restore them: we explain what to ask the supplier for and stop
  • The importer is a closed product whose code you cannot change: we list the settings to try instead

Questions

Can you repair names that are already garbled in my database?

Not in this job. If the original file is still available, reconciliation is a separate agreed scope. If only the garbled copy remains, some characters cannot be recovered.

Which encoding will the importer use?

The one agreed for your named supplier layout, checked against your sample. We do not make the importer guess for unknown files.

Why not rely on the default encoding of the importer's language?

Defaults differ by language, version and platform. Python's csv documentation, for example, says a file opened for reading uses the system default encoding in Python 3.14 and UTF-8 in Python 3.15. So the importer names the agreed encoding explicitly rather than relying on a default.

Does it need my live supplier feed?

No. We work from a few invented or redacted rows. A live download is not needed and should not be sent.

Price and terms

£195 · untested offer price. £195 after the agreed test file imports with every agreed value unchanged and you sign off. No payment before sign-off.

This is a new service with no published client results. The price is a starting point we have not yet tested with buyers. Nothing is ordered or charged by the enquiry. The full specification is on the Synthetic Industry catalogue.

Request this outcome

Enquiry about: Make a supplier CSV import keep accented and special characters intact. Page: /services/csv-encoding-bom-garbled-characters-import/.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

An enquiry is not an order. We assess fit and agree scope, safe access and terms before any work. Prefer email? hello@syntheticindustry.ai.