Biology Unbound
FeaturesLong read

Linear DNA Templates for Rapid CFPS Variant Generation

PCR-based linear templates compress protein variant testing from days to hours.

Staff Writer · · 13 min read
Cover illustration for “Linear DNA Templates for Rapid CFPS Variant Generation”
Features · August 28, 2026 · 13 min read · 2,945 words

The gap between a protein variant idea and an expressed, testable protein used to run three days deep: clone the gene, transform it into cells, grow the culture, prep the plasmid. Linear DNA templates cut that same path to about three hours, using PCR instead of cloning. Cell-free protein synthesis is the only platform where that speed advantage actually survives contact with a real experiment, and this piece explains why.

At small scale, three days per construct is annoying, not fatal. But protein engineering campaigns rarely stop at one construct, and modern design cycles call for dozens, sometimes hundreds, of variants tested in parallel. When each one carries a three-day cloning tax, cloning becomes the bottleneck for the whole pipeline, no matter how fast the assay or how good the model picking variants. A PCR reaction that produces a usable linear expression template (LET) in three hours changes that math completely. What has to be true, at the level of chemistry and reagent design, for that three-hour promise to actually hold up?

What makes CFPS uniquely compatible with linear DNA

Cell-based expression has a hard requirement most people outside molecular biology never think about: the DNA has to be circular. Stable episomal expression inside a living cell depends on supercoiled plasmid, because linear DNA introduced into a cell gets flagged, degraded, or just fails to replicate across generations. There's no way around this if the goal is a stable transformed cell line, since selection pressure, replication origin function, and antibiotic resistance maintenance all assume a circular template.

Cell-free protein synthesis runs under none of those constraints. There are no generations to keep stable across, no transformation step, no living cell that needs to propagate the DNA at all. A PCR product goes straight into the reaction tube as the template, and the system carries out transcription and translation from whatever DNA sits in the mix. This coupled transcription-translation format, both processes happening in one reaction vessel, is now the dominant setup in CFPS systems, and it's the reason direct DNA-to-protein conversion is possible without a cloning step in between.

CFPS gets called an "open system" often enough that the phrase risks losing its meaning, but here it earns its keep: reaction conditions are directly accessible and tunable in a way the inside of a living cell never is. That openness isn't incidental. It's why the template itself can be supplemented, modified, or protected using tools that simply don't work inside a cell membrane, and it sets up everything that follows about nuclease protection.

The nuclease problem that kept LETs from being practical

The same crude extracts that make CFPS work also carry the extract's original nucleases along for the ride. In E. coli-derived systems, the main offender is RecBCD, a helicase-exonuclease complex that chews through linear DNA with a processivity that makes it very good at its native job and very bad news for anyone trying to use a PCR product as a template. Left unaddressed, RecBCD activity means LET-based reactions produce much lower protein yields than the same construct delivered as plasmid, and sometimes fail to produce a usable amount of protein at all.

This isn't an E. coli quirk, either, since V. natriegens extracts, which have drawn interest for their fast growth rate and use in CFPS, show similar exonuclease activity against linear templates. Degradation is a general property of crude lysate systems, not a peculiarity of one organism's biology.

For years, this is exactly why LETs stayed a nice idea on paper rather than a working tool. The time saved at the template prep stage got eaten right back up by yield losses in the reaction itself; a three-hour PCR product that expresses at a fraction of plasmid yield isn't actually faster, it's just differently slow. Most of the recent progress in CFPS engineering has gone toward solving exactly this problem, and the protection methods now available are what turned LETs from a theoretical shortcut into something a working lab can put on a repeatable protocol sheet.

Protection strategies that make LET yields competitive with plasmid

Three distinct approaches have emerged to deal with the nuclease problem, and they intervene at different points, asking different things of the person running the reaction.

The first is GamS protein supplementation. GamS comes from a bacteriophage and works by physically blocking the active site of Exonuclease V, the RecBCD complex, so the enzyme can't engage the DNA anymore. It's borrowed defense, really: bacteriophage DNA faces the same threat from RecBCD that a linear template does, and GamS is the phage's own evolved countermeasure, repurposed here as an additive. It drops into an existing extract with no need to re-engineer the strain the extract came from, which makes it the lowest-friction option for a lab that didn't build its own extract pipeline.

The second is terminal DNA-binding protein protection, sometimes called the CroP-LET method. Instead of inhibiting the nuclease, this approach attaches a DNA-binding protein (based on the scCro protein) directly to the ends of the linear template, capping the exposed DNA ends where exonucleases start chewing. In E. coli CFPS, this method raised sfGFP expression from LETs sixfold over unprotected template. In V. natriegens CFPS, where the nuclease environment is apparently more aggressive, the same capping strategy pushed sfGFP yields up by as much as 18-fold, reaching 0.3 mg/mL. Sit with that scaling relationship for a second: the harsher the extract's native nuclease activity, the more this method has to offer, which suggests the mechanism does exactly what it claims, blocking physical access rather than working around the problem some other way. It's modular, too, because the protection gets added at the protein level rather than baked into the DNA sequence, so it can in principle apply to any LET without touching the template design.

The third is extract engineering: building the source strain without the nuclease genes in the first place. Deleting recBCD from the organism used to make the extract removes the threat at the source rather than suppressing it after the fact in every single reaction. This costs more up front, since it means investing in strain construction rather than just buying an additive, but it means never having to remember to add an inhibitor to every tube. It's a permanent fix built into the reagent itself, not a per-reaction workaround.

None of these three wins outright over the other two. They're a tiered toolkit, and which one fits a given lab depends on what extract they're using, how much throughput they need, and whether they control their own extract manufacturing or buy it from someone else. Taken together, though, they close the yield gap that used to make LETs a trade-off: with adequate protection, LET expression now performs comparably to plasmid-based expression, which is the condition that has to be met before the speed advantage means anything at all.

Designing LETs that work: sequence architecture from promoter to terminus

A linear expression template isn't just a PCR amplicon of a coding sequence. It needs the full set of regulatory parts arranged in the right order and spacing for in vitro transcription and translation to actually happen. At minimum: a promoter the extract's RNA polymerase recognizes (a T7 promoter for T7-based systems, or an endogenous sigma factor promoter for extract-coupled systems), a ribosome binding site positioned at the correct distance from the start codon, the coding sequence itself including any tag or fusion partner, and a transcription terminator marking the 3' end of the message.

These pieces get stitched together with overlap-extension PCR or Gibson assembly, using standardized primers at the outer edges. Here's the part that matters for anyone running a variant library rather than a single construct: because the promoter, RBS, and terminator are identical across every variant in the library, one pair of universal outer primers can amplify the whole set off a common backbone. Only the coding sequence changes from variant to variant, so primer design and primer cost stop scaling with library size, and that matters a great deal once a library runs into the dozens or hundreds.

Assembly itself can push to plate scale. A cell-free method combining Gibson assembly with direct PCR amplification off the unpurified assembly product, run entirely in 384-well plates, can produce a usable LET in under three hours, with no cell culture step anywhere in the process. Once template construction happens on the same plate format as the expression reaction, DNA design and protein expression stop being two separate stages of work with a wait in between, and they become one continuous plate-to-plate workflow.

How a variant screening run actually works end to end

An ML-directed protease evolution workflow, published in ACS Synthetic Biology in 2025, gives the clearest documented picture of what LET-CFPS screening looks like at real, working scale, not as a proof-of-concept.

The workflow runs in three phases. Create: 48 random variants covering a target region get encoded as linear DNA and combined directly with CFPS reagents, no cloning involved. Test: the protease variants are expressed and assayed in the same plate, using a FRET substrate whose fluorescence rises as protease activity increases, giving a real-time readout of function. Learn: the activity data trains a machine learning model that identifies which sequence positions actually matter, and that model designs the next round of variants.

Across the full campaign, the study screened 138 variants total, with each full production-and-screening cycle completing in six hours. The team started with 48 random variants to sample the fitness landscape broadly, followed up with 32 targeted variants chosen by the ML model, and came out the other end with protease variants showing improved kinetic properties. Six hours per cycle means a team can run several complete design-build-test-learn loops in a single week, something that's simply out of reach when each loop needs multi-day cloning before testing can even start.

Worth flagging: the FRET assay in this workflow reads out in the same plate where the protein was expressed, with no transfer step, no purification, no separate handling of the protein between expression and assay. That's only possible because CFPS produces protein in a format that's already open and assay-ready, rather than locked inside a cell that needs lysing first. A related approach pushes this same plate-based logic further by adding robotic liquid handling, and has been used to screen thousands of enzyme variants for retained activity after thermal challenge, again via a fluorescence readout in the same plate. Same architecture; robotics just adds scale on top of it.

Proteins that gain the most from moving variant work into CFPS

Every protein benefits from faster iteration, in principle. But for some protein classes, CFPS isn't just faster, it's the only route that produces usable data at all.

Toxic proteins are the clearest case. A protein that kills or impairs its host cell on overexpression produces poor yields in any cell-based system, and typically forces researchers into inducible expression systems with a narrow window of viable induction before the culture crashes. CFPS has no host cell to protect, so there's no viability ceiling limiting expression in the first place.

Proteins prone to forming inclusion bodies face a related problem: overproduction in cells frequently drives misfolded protein into insoluble aggregates. CFPS reactions let a researcher directly tune redox potential, ionic strength, and chaperone concentration in the tube, adjusting the folding environment in ways that just aren't accessible once a protein is being made inside an intact cell.

Disulfide-rich proteins present their own challenge, since the reducing environment of the bacterial cytoplasm actively fights disulfide bond formation. In CFPS, pretreating the extract with iodoacetamide inactivates the cytosolic redox enzymes that would otherwise undo disulfide pairing, and adding an oxidized/reduced glutathione buffer along with the chaperone DsbC lets correct disulfide bonds form during synthesis. This combination has produced proteins with as many as 24 disulfide bonds, at large scale, a class of target that standard E. coli cell-based expression handles poorly if at all.

Membrane proteins get a boost from a different feature entirely: wheat germ extract systems carry native microsomes that can solubilize membrane proteins without adding exogenous liposomes, an advantage specific to eukaryotic extract-based CFPS.

Multi-domain and glycosylated proteins round out the list. Eukaryotic CFPS systems built around native microsomes have produced correctly formed disulfide bonds and high-mannose N-glycan patterns for targets that include a monoclonal antibody, the SARS-CoV-2 receptor-binding domain, and a G protein-coupled receptor, with mass spectrometry used to confirm results in some cases.

For variant engineering across any of these classes, cell-based expression doesn't just run slower. It adds noise: toxicity effects, inconsistent inclusion body formation, host-cell proteolysis, the kind of static that muddies the signal a researcher is actually trying to read off the variant itself. CFPS removes that noise rather than just working around it.

Where cell-based expression still has the advantage

None of this makes CFPS a universal replacement, and nobody serious claims it is. CFPS reactions are time-limited: a typical reaction stays productive for hours, not the days or weeks a fed-batch cell culture can sustain, and batch yield per reaction volume is constrained accordingly.

For proteins that need complex eukaryotic post-translational modification, particularly site-specific glycosylation matching a therapeutic-grade standard, cell-based mammalian expression remains the more mature, better-validated platform. Eukaryotic CFPS systems haven't reached manufacturing scale yet, and bacterial CFPS scales more easily but can't deliver eukaryotic modifications, so neither side of the CFPS world closes that particular gap on its own.

Lot-to-lot consistency in extract preparation is a real limitation too, since extract activity varies from batch to batch, and without careful QC documentation at the individual lot level, reproducibility across a long screening campaign gets shaky fast.

The practical division of labor follows pretty directly from all this: use LET-CFPS to screen variant libraries fast and narrow down to high-confidence leads, then hand those confirmed leads off to cell-based production for scale-up, for post-translational modification fidelity, or for regulatory-grade manufacturing. CFPS isn't competing with cell-based systems on their home turf. It's the front end of the pipeline, the stage where fast decisions get made before anyone commits the time and cost of downstream cell-based work.

What reproducible LET-CFPS screening actually requires from the reagent system

A variant screening campaign is, at its core, a comparison exercise: variant A beat variant B because it showed higher activity in the assay. That comparison only means something if the expression environment was actually equivalent across every well and every day the campaign ran.

Nuclease activity isn't identical from one extract lot to the next, and that variation feeds straight into LET stability and yield. An extract lot with higher residual exonuclease activity degrades LETs faster than a cleaner lot, which compresses the signal from lower-expressing variants and can make a genuinely weak variant look worse than it is, or a mediocre one look artificially competitive.

This is why lot-level QC documentation matters as much as it does: a manufacturer reporting actual measured nuclease suppression performance, yield benchmarks, and activity data for each lot gives a researcher what they need to know whether two experiments run weeks apart are actually comparable, or whether an apparent difference in variant performance is just reagent drift. Formulation transparency plays the same role. If a researcher can't see what's actually in the reaction, there's no way to tell whether a yield drop in a new lot came from extract quality, LET protection efficiency, or something else altogether; a black-box reagent just dumps that whole troubleshooting burden onto the person running the experiment.

Automation adds another layer of demand. Running LET-CFPS at 384-well throughput with robotic liquid handling requires that reagent viscosity, dispensing behavior, and reaction kinetics stay consistent lot to lot, not just that the average yield looks similar on paper. Variability invisible in a hand-pipetted, 8-tube experiment turns into a systematic error once it's repeated across hundreds of wells.

OpenCFPS™ from Sepia Biosciences handles this by publishing lot-level QC data and documenting its formulations rather than keeping them proprietary, which is exactly the kind of transparency that closes the reproducibility gap described above. It's also priced to make large-scale screening workable, rather than forcing a lab to cut replicate counts just to stay within budget.

Building a LET-CFPS variant pipeline: practical decisions for a working lab

Putting a working LET-CFPS pipeline together comes down to a handful of decisions made in the right order, not one technology purchase. Start with template design: build a universal outer-primer system around a fixed promoter, RBS, and terminator so only the coding sequence varies across the library, which keeps primer cost and design time from scaling with library size.

Next, pick a nuclease protection strategy that matches the extract in hand: GamS supplementation for a lab working with an existing extract source, terminal capping for a lab that wants a modular fix independent of extract origin, or investment in a nuclease-deficient strain for a lab producing its own extract at volume. After that comes committing to plate-based assay integration from the start, so expression and readout happen in the same well rather than adding a transfer or purification step that reintroduces the delay LETs were supposed to eliminate.

The piece most often shortchanged is reagent QC discipline: insist on lot-level activity data and documented formulations from whatever CFPS system a lab buys, because without that visibility, a multi-week screening campaign risks mistaking reagent variation for genuine biological signal. None of this is exotic or expensive on its own, but together, it's what decides whether the three-hour promise of a linear template turns into a working pipeline, or just a faster way to generate DNA that still sits waiting for a reaction that can't make good use of it.

Sources

  1. pubs.acs.org
  2. pubs.acs.org
  3. ncbi.nlm.nih.gov

More in Features