Index
Year
2026
Field
Computational biology
Status
Ongoing
Project
1 of 4

Proteins against microplastics

I validated a docking pipeline, then ran it against a protein whose experimental answer key already existed. It scored all six known non-binders as credible binders.

Docking scorekcal/mol · lower binds tighter

  1. TerephthalateThe real ligand−8.06Binds · Kd 0.364 µM
  2. HomophthalateRegioisomer-type−7.48No binding
  3. HippurateMonocarboxylate−7.10No binding
  4. PABAMonocarboxylate−6.84No binding
  5. BenzoateMonocarboxylate−6.75No binding
  6. TyrosineMonocarboxylate · amino acid−6.51No binding
  7. PhenylalanineMonocarboxylate · amino acid−6.04No binding

Gautom et al. 2021: sixty-one compounds screened

TphC closed around terephthalate (orange), from PDB 7NDS. Rust marks the hinges.

01 · The question

A clamshell for plastic

Microplastics are silently causing health issues and no one seems to be doing much about it. I decided to start with a popular plastic, PET, that breaks down into terephthalate. A bacterial protein called TphC binds it and encases it like a clamshell, which makes it a start for a sensor that detects plastic being broken down. With this, I set out to design a variant with sharper selectivity.

02 · The pipeline

It found the crystal pose

I built a docking pipeline against the closed crystal structure, PDB 7NDS, and validated it first: redocking the native ligand reproduced the crystal pose to 0.28 Å.

03 · The designs

Thirty variants, none better

Then I generated thirty designed variants, changing three positions in the pocket, and ranked them against a panel of human metabolites that might compete for the site. None of the thirty beat the wild-type protein. I assumed my designs were bad.

0 of 30 beat the wild type

04 · The screen

Six compounds, six credible binders

The panel was six compounds a sensor would have to ignore: five metabolites it would meet in a cell, and one close structural cousin of terephthalate. My pipeline docked each into the pocket, and scored all six as credible binders.

05 · The answer key

The paper had already tested them

Then I read the results section of the paper that deposited the structure. Gautom and colleagues had screened sixty-one compounds. Terephthalate binds at a Kd of 0.364 µM; none of the six binds measurably. They are monocarboxylates and regioisomers, the classes that paper reports as non-binding.

06 · Why

The score can’t see what’s missing

On the real ligand the prediction was good. On the non-binders it was wrong by two to four kcal/mol, always in the same direction. A scoring function adds up the contacts it finds, and has no term for a group the pocket requires and the ligand lacks: a benzene ring and one carboxylate collect most of terephthalate’s score, without the second carboxylate the pocket demands.

07 · The range

2,389-fold, squeezed to 15%

Against six ligands spanning 2,389-fold in measured affinity, the scores span 15% of that range. A re-parameterised scoring function gives 17%, and errs further on the real ligand, so swapping functions is not a fix.

−9−8−7−6−5−4MeasuredDocking score

4.6 kcal/mol measured · 0.7 scored

08 · A second result

Confident, in the wrong state

ESMFold predicts this protein accurately, in the wrong state. It reproduces the open, empty structure to 0.92–1.09 Å across all six chains, while missing the closed one by 4.014 Å, at a mean confidence of 91 out of 100. Nothing in the output says which state you have been given.

09 · Your turn

Run the screen yourself

Dock a compound, read its score, and call it.

Pick a compound to dock it into the pocket.

10 · What I would do differently

Check the answer key first

Validating pose prediction and then relying on affinity ranking are two different tasks, with two different error profiles. The check that would have caught it was the published screen: one search away, and four years old.

The rule I took from it is narrower: assert on output properties, never on exit codes.

The full write-up follows

The question

Microplastics are silently causing health issues and no one seems to be doing much about it. I decided to start with a popular plastic, PET, that breaks down into terephthalate. A bacterial protein called TphC binds it and encases it like a clamshell, which makes it a start for a sensor that detects plastic being broken down. With this, I set out to design a variant with sharper selectivity.

What I did

I built a docking pipeline against the closed crystal structure, PDB 7NDS, and validated it: redocking the native ligand reproduced the crystal pose to 0.28 Å. Then I generated thirty designed variants and ranked them against a panel of human metabolites that might compete for the site.

None of the thirty beat the wild-type protein. I assumed my designs were bad.

What I found

Then I read the results section of the paper that deposited the structure I had been docking into. Gautom and colleagues had already screened sixty-one compounds by differential scanning fluorimetry, with calorimetry on the hits. Wild-type TphC binds terephthalate at a Kd of 0.364 µM, and does not measurably bind any of the six compounds I had spent weeks designing against. They are monocarboxylates and regioisomers — the classes that paper reports as non-binding.

My pipeline had scored all six as credible binders.

Predicted binding scores for seven compounds. Terephthalic acid scores −8.06 and binds in experiment, at a Kd of 0.364 µM. The other six score between −7.48 and −6.04, close enough to be read as binders, and none of them binds in experiment.

Test Result
Error on the real ligand +0.73 kcal/mol — accurate
Non-binders scored as credible binders 6 of 6
True selectivity margin understated by 2.8–4.1 kcal/mol
Designed variants beating wild type 0 of 30

The design campaign was not failing. It was correctly answering a question that should never have been asked.

The error is asymmetric, and that is the interesting part. On the real ligand the prediction was good. On the non-binders it was wrong by two to four kcal/mol, always in the same direction. The reason looks structural. An empirical scoring function adds up steric, hydrophobic and hydrogen-bond terms over whichever atoms are in contact. It has no term for this ligand is missing a group the pocket requires. A benzene ring plus one carboxylate collects most of terephthalate’s contact score, even though the second carboxylate the pocket actually demands is absent. Both of the aliphatic controls were rejected correctly, which locates the failure precisely: the function tells gross shape apart, not specific polar recognition.

Against six ligands spanning 2,389-fold in measured affinity, the scores span 15% of that range. A re-parameterised scoring function gives 17%, and errs further on the real ligand, so swapping functions is not a fix. Re-docking into three chains of the open structure shows this is not an artefact of using a pocket that had closed around the ligand.

A second result, from the same two structures

The model above can be pulled between the two crystal structures: closed around terephthalate, and open and empty. That motion is also where the project’s other finding came from.

ESMFold predicts this protein accurately — in the wrong state. It reproduces the open, empty structure to 0.92–1.09 Å across all six chains, while missing the closed one by 4.014 Å. Mean confidence was 91 out of 100 throughout, and confidence correlates with per-residue error at r = −0.26. Nothing in the output says which state you have been given, and for a protein that only binds when closed, that is the difference between a usable receptor and an unusable one.

What I would do differently

Validating pose prediction and then relying on affinity ranking are two different tasks with two different error profiles. I had flagged that limitation in my own notes, in writing, and then built three weeks of work on the output anyway. The check that would have caught it was the published screen — one search away, and four years old.

The rule I took from the tooling side is narrower and has been more useful since: assert on output properties, never on exit codes. Five separate failures in this pipeline produced plausible numbers rather than errors. One of them wrote a zero-byte receptor file, and the docking program then reported a successful run against an empty protein.

Honest limitations

None of the three findings is novel, and I checked: scoring functions tracking molecular size, scoring functions disagreeing with each other, and structure predictors returning the empty state are all documented. What is unusual here is the worked example — a pipeline scored against the experimental panel from the very paper that produced the structure, with the failure quantified end to end and the code public. The affinity benchmark rests on six ligands, which is too few for the correlations to mean much; the false-positive count is the part that survives the sample size. Nothing here has been tested experimentally.

Where it goes next

The write-up is a complete draft aimed at bioRxiv, not yet submitted. The protein work led somewhere else: whether laboratory demonstrations of plastic-degrading enzymes inside cells survive contact with real plastic, which is more crystalline than the material those demonstrations use. That question is open, and it needs a sensor that works in a mammalian cell, which does not yet exist.