Taylor McNeil

docs-as-portfolio v1.9

PUT/aampersand/an-alphabet-too-short

An Alphabet Too Short

How I built a manuscript importer for Scrivener, Word, and Markdown in 63 days, and why perfect import fidelity was never the goal.

Context

This is the design companion to Tilted Hermai. That piece is about crossing the distance between two languages. This one is about what's actually inside the files we toil in.

Follow the Cow

I copied and pasted my chapters into aampersand for months. Like a Luddite.

It worked for me. It was simple. I was writing new chapters directly in aampersand, so the copy-paste tango wasn't one I did often. Things were fine.

Then I looked at the calendar: 63 days until friends and family. Until hands that weren't my own were going to touch aampersand, and it didn't have an importer.

Oops.

For most apps, import is a file upload. Pick a file, it lands somewhere, done. A manuscript presents a problem for aampersand. A manuscript presents a problem for aampersand, because it doesn't really store your book as one giant text blob. It deconstructs your book into chapters, scenes, and prose. A normal upload is one row. A manuscript is hundreds. I can't just take your Markdown or your Word doc and dump it on a page. Something has to read it, figure out where your chapters are, and rebuild your book in the aampersand shape.

The real question, at 63 days and counting: who was going to build that? (Me?)

So, like all great 10x cracked-out developers (which I am not), I started with the hardest problem first.

Scrivener.

My thought process was this:

  • Scrivener users love Scrivener. If you can convert them, you can convert anyone. They live and breathe by their folders, their corkboards, and research binders.
  • Scrivener has its own file format, and it's the weirdest one.
  • If I do the hardest one first, everything downstream will be easy.

Two of those turned out to be true.

Letters from Phoenicia

My naive assumption was that "structure" meant roughly the same thing everywhere. Markdown has # for headings. Word would have something like that. Scrivener too. Build one flexible template for "a book," point three parsers at it, done. I just needed to sit down and open a few files and easy money. Bob's your uncle.

That is not how any of this works.

Every format has a completely different idea of what a chapter is, and some of them don't have an idea at all. So before I could translate anything, I had to crack each one open and see what was actually inside.

Scrivener: the filing cabinet

Lesson #1: The hard road is empty for a reason.

A .scriv project isn't a document. It's a folder of folders. Zip it up and you're handing me a small filesystem.

Loading diagram...
Annotated diagram

The binder is a treasure trove. It's a curated table of contents from the writer. This folder is Part One. These documents are chapters. That one is research. That one is trash (don't import that). I'm not guessing at anything. I'm reading intent the writer declared.

The adapter opens the zip, reads the one .scrivx manifest, walks the binder tree, and for every document pulls the prose, the synopsis, and the notes. The prose is RTF (I mean I get it, but also...😑). Tables, diagrams, and research references get flagged for later.

Scrivener was hard, but not for the reason I expected. The prose was fine. Everything else was the problem, and I made it worse on purpose. My first version skipped the research files entirely, and it felt too easy. Some might say I let my head get quite big.

So I decided that if Scrivener hands me a research folder, its contents should land somewhere. That decision is how aampersand got a Vault. (This was a wonderful decision that did not at all dramatically alter the product, add tons of extra feature work, and require massive amounts of coordination on my part for a singular decision...not at all...)

Word: the wild west

Here's a fun fact. A .docx file is actually a zip file. It contains way more than just text. It contains prose, styles, and comments. At its heart is one long XML document with styles layered on top.

But therein lies the problem, dear reader.

Because you can give anything meaning in Word. Heading 2 could be your chapter headings. Heading 3 could be the title heading. You can do all kinds of wacky things in Word. YOU CAN PUT TABLES IN THE PROSE. (I'M LOOKING AT YOU GAMELIT WRITERS!)

Loading diagram...
Annotated diagram

The adapter reads the XML, converts it to HTML, then starts inferring. One H1 in the whole file? Probably the title. Twenty-two of them? Probably chapters. H2s might be scenes, or they might be subheadings inside a chapter. Comments and explicit highlights get pulled out so they don't end up baked into the prose.

Word doesn't tell you what the writer meant. It tells you how the writer formatted, and you work backward.

Markdown: the golden child

I should have started here.

Loading diagram...

That's the whole diagram. Not a lot to the format, to be honest.

No metadata. No synopsis. No notes. Headings and horizontal rules are the only structural signals you get. The adapter decodes the text, scans for headings and scene breaks, and infers parts, chapters, and scenes from how the headings nest relative to each other.

Three formats, three levels of confidence

Line them up next to each other, and I think you can see the problem.

StructureSynopsisNotesResearch
ScrivenerExplicit✓✓✓
WordInferred———
MarkdownVibes———

So the one-template idea died. Each format gets its own parser, speaking its own dialect, and each one does its own translation into the same thing on the other side: a proposed project tree of parts, chapters, and scenes, with prose converted to HTML.

Sowing Dragon Teeth

Parsing is the easy part. What comes out of it is a pile of guesses, and some of those guesses are going to be wrong.

Here's the full shape of the pipeline, once all three lanes are done doing their specialist work:

Loading diagram...

Show your work

The most important decision in the whole pipeline is that nothing goes straight into the database. Every import stops at a review surface first, and the posture of that surface is always the same: here's what we think is going on. Tell us where we're wrong.

I would show you a picture of the original jank screen, but I just redesigned it, and it looks much better. I will spare your eyes.

The best part of the system is that the writer can overrule everything. This isn't a part. This isn't a chapter. This is a scene break. Don't bring this in. The system proposes. The writer decides.

The table problem

Remember those GameLit writers? They came back to haunt me. Tables showed up in all three formats, forcing me to make a decision about them.

Like any good researcher, I took to the streets. What did the public have to say about importing and tables? To be honest, it was a mixed bag.

AppScrivener importTables in the manuscript editorOther imports
NoveliumYes, native .scrivxCouldn't verifyDOCX, TXT, PDF, Google Docs
NovellumNoNoDOCX, Markdown, TXT
DabbleNoProperty lists, not tablesDOCX, Markdown, HTML, RTF
NovelcrafterCouldn't find a native importerCouldn't verifyDocument import workflows
CampfireNoStat panels in worldbuilding, not manuscriptEPUB, DOCX, .blaze
PlottrYes, .scriv projectsN/A, not a prose editorDOCX, Snowflake Pro
Scrivener—YesDOCX, RTF, ODT, TXT, Markdown, and more

I spent maybe twenty minutes on the problem and decided, nah. Tables suck in any tool, including Word. The editor is for prose, and specifically prose that's going to become a book. Outside of LitRPG stat blocks, nobody is putting tables in a novel headed for an EPUB.

But I'm also not going to throw away your stuff (kind of me, I know). So tables get converted to Markdown and routed to the Vault. You still have them. You can reference them. They just don't get to pretend they're prose. (May come back and rethink this later, but I am only one person, and this app needs to get off my computer.)

Seven Gates is Plenty

Around the third week of edge cases, I had a very specific realization.

I could make this better forever. There are people who only write scenes and never make chapters. People who separate scenes with three ampersands. People who use emojis. People whose Scrivener binders are...unhinged. I will never think of all of them. Nobody will.

Say it takes 300 hours to chase every possible edge case. Every one of those hours serves a smaller slice of writers than the last. And import is something you do once. Maybe twice. You walk through the door and you don't think about the door again.

Lesson #2: A promise of perfection is a commitment to lie.

Perfect fidelity was never the goal. No tool promises it. Nobody expects it.

What writers expect is that you get the big things right, you don't lose their work, and you let them fix the rest in five minutes. That's the promise worth making.

This summer, aampersand learned to read other people's books. In September, we were swallowed whole.

Next month: an Olympic sized headache.

aampersand

Bring your book with you. aampersand shows you what it found and lets you decide the rest.

Read: Tilted Hermai