Skip to content

Commit 9f33212

Browse files
author
benquist
committed
Fix 2021 publications sync
1 parent 6ea585d commit 9f33212

7 files changed

Lines changed: 397 additions & 107 deletions

File tree

.github/workflows/sync-google-doc-cv.yml

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -22,15 +22,17 @@ jobs:
2222
python-version: "3.12"
2323

2424
- name: Sync CV from Google Docs export
25-
env:
26-
GOOGLE_DOC_CV_ID: ${{ secrets.GOOGLE_DOC_CV_ID }}
2725
run: |
2826
bash scripts/sync_google_doc_cv.sh
2927
3028
- name: Sync new publications from CV DOIs
3129
run: |
3230
python scripts/sync_publications_from_cv.py
3331
32+
- name: Rebuild publications include from public Google Doc HTML
33+
run: |
34+
python scripts/rebuild_publications_include_from_doc.py
35+
3436
- name: Sync publications HTML from CV
3537
run: |
3638
python scripts/sync_publications_html.py
@@ -52,6 +54,7 @@ jobs:
5254
assets/cv/publications_sync_report.txt \
5355
_data/cv_sync.yml \
5456
_bibliography/papers.bib \
57+
scripts/rebuild_publications_include_from_doc.py \
5558
_includes/publications_full_from_doc.md
5659
git commit -m "chore(cv): sync CV and publications from Google Doc"
5760
git push

_includes/publications_full_from_doc.md

Lines changed: 34 additions & 31 deletions
Large diffs are not rendered by default.

_pages/publications.html

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -581,6 +581,10 @@ <h2>Complete Publication List</h2>
581581
label: 'Conservation Impacts',
582582
threshold: 2,
583583
matchers: [
584+
{ pattern: /deregulation/i, weight: 2 },
585+
{ pattern: /increasing fire|landscape fire|wildfire/i, weight: 2 },
586+
{ pattern: /amazonian biodiversity/i, weight: 3 },
587+
{ pattern: /drought.{0,30}biodiversity|biodiversity.{0,30}drought/i, weight: 2 },
584588
{ pattern: /conservation planning/i, weight: 3 },
585589
{ pattern: /conserving terrestrial biodiversity/i, weight: 3 },
586590
{ pattern: /biodiversity, carbon and water/i, weight: 3 },

assets/cv/publications_sync_report.txt

Lines changed: 3 additions & 48 deletions
Original file line numberDiff line numberDiff line change
@@ -1,57 +1,12 @@
1-
Publications Sync Report — 2026-06-06 06:54 UTC
1+
Publications Sync Report — 2026-06-06 07:56 UTC
22
============================================================
33

44
CV entries parsed: 332
55
In Press → published patches applied: 0
6+
Missing-paper insertions applied: 0
67
DOIs verified OK: 23
78
DOI failures / anomalies: 3
8-
Papers in CV but NOT found in HTML: 43
9-
10-
MISSING FROM WEBSITE HTML (in CV but not found in _includes/publications_full_from_doc.md)
11-
----------------------------------------
12-
[2025 #11] Gaudard, J., Telford, R. J., Chacon‐Labella, J., Dawson, H. R., Enquist, B. J., Töpper, J. P., ... & Halbritter, A. H. (2025). fluxible:An R …
13-
[2024 #27] Westgeest, Adrianus J., François Vasseur, Brian J. Enquist, Rubén Milla, Alicia Gómez‐Fernández, David Pot, Denis Vile, and Cyrille Violle. …
14-
[2023 #50] Jónsdóttir, Ingibjörg S., et al. "Intraspecific trait variability is a key feature underlying high Arctic plant community resistance to clim …
15-
[2022 #61] Chaplin-Kramer, R., Brauman, K. A., Cavender-Bares, J., Díaz, S., Duarte, G. T., Enquist, B. J., ... & Zafra-Calvo, N. (2022). Conservation …
16-
[2022 #62] Feng, X., Enquist, B. J., Park, D. S., Boyle, B., Breshears, D. D., Gallagher, R. V., ... & López‐Hoffman, L. (2022) A review of the hetero …
17-
[2022 #63] Boyle, B. L., Maitner, B. S., Barbosa, G. G., Sajja, R. K., Feng, X., Merow, C., ... & Enquist, B. J. (2022). Geographic name resolution ser …
18-
[2022 #64] Maitner, B. S., Park, D. S., Enquist, B. J., & Dlugosch, K. M. (2022). Where we've been and where we're going: the importance of source comm …
19-
[2022 #65] Aguirre-Gutiérrez, Jesús, et al. Functional susceptibility of tropical forests to climate change. Nature Ecology & Evolution (2022): 1-12. …
20-
[2022 #66] Wieczynski, D.J., Díaz, S., Durán, S.M., Fyllas, N.M., Salinas, N., Martin, R.E., Shenkin, A., Silman, M.R., Asner, G.P., Bentley, L.P., Mal …
21-
[2022 #67] Lourenço Junior J, Enquist BJ, von Arx G, Sonsin Oliveira J, Morino K, Thomaz LD, Milanez CR. (2022) Hydraulic tradeoffs underlie local var …
22-
[2022 #68] Gatti, R.G. et al. (2022) The number of tree species on Earth. Proceedings of the National Academy of Sciences, U.S.A DOI: 10.1073/pnas.211 …
23-
[2022 #69] Maitner BS, Park DS, Enquist BJ, Dlugosch KM. (2022) Both source and recipient range phylogenetic community structure can predict the outcom …
24-
[2022 #70] Brun, P., Violle, C., Mouillot, D., Mouquet, N., Enquist, B.J., Munoz, F., Münkemüller, T., Ostling, A., Zimmermann, N.E. and W. Thuiller ( …
25-
[2021 #71] Enquist, B. J. (2021) How Pandemics Rapidly Reshape the Evolutionary & Ecological Landscape, In: The Complex Alternative: Complexity Scient …
26-
[2021 #72] Feng X, Merow C, Liu Z, Park DS, Roehrdanz PR, Maitner B, Newman EA, Boyle BL, Lien A, Burger JR, Pires MM. P. M. Brando, M.B. Bush, C.N.H. …
27-
[2021 #74] Huang, C.Y., Durán, S.M., Hu, K.T., Li, H.J., Swenson, N.G. and B.J. Enquist (2021) Remotely sensed assessment of increasing chronic and epi …
28-
[2021 #75] Chaplin-Kramer, R., Brauman, K.A., Cavender-Bares, J., Díaz, S., Duarte, G.T., Enquist, B.J., Garibaldi, L.A., Geldmann, J., Halpern, B.S., …
29-
[2020 #93] Vandvik, V., Halbritter, A.H., Yang, Y., He, H., Zhang, L., Brummer, A.B., Klanderud, K., Maitner, B.S., Michaletz, S.T., Sun, X. and Telfor …
30-
[2019 #104] Durán, S.M., Martin,R.E., Díaz, S., Maitner, B.S., Malhi, Y., Salinas, N., Shenkin, A., Silman, M.L., Wieczynski, D.J., Asner, G.P., Bentl …
31-
[2019 #112] Weiser, M.D., Ning. D., Buzzard, V., Michaletz, S.T., He., Z., Enquist, B.J., Waide, R.B., Zhou, J., and M. Kaspari. Thermal disruption of s …
32-
[2019 #118] Šímová, I., Sandel, B., Enquist, B.J., Michaletz, S.T., Kattge, J., Violle, C., McGill, B.J., Blonder, B., Engemann, K., Peet, R.K. and Wise …
33-
[2018 #120] Bjorkman, A.D. et al. (2018) (50+ international authors) Plant functional trait change across a warming tundra biome Nature, 562:57–62. …
34-
[2017 #146] Blonder, B., Moulton, D.E., Blois, J., Enquist, B.J., Graae, B.J., Macias-Fauria, M., McGill, B.J., Nogué, S., Ordonez, A., Sandel, B., and …
35-
[2016 #164] Feakins, S.J, T. Peters, M.S. Wu, A. Shenkin, N. Salinas, C.A.J. Girardin, L. P. Bentley, B. Blonder, B. J. Enquist, R.E. Martin, G. P. Asne …
36-
[2016 #173] Enquist, B.J. (2016) Biology Distilled. Nature, 531:34. …
37-
[2015 #179] Duncanson, L.I., Dubayah, R.O, and B. J. Enquist Assessing the general patterns of forest structure: Quantifying tree and forest allometri …
38-
[2015 #180] Blonder, B., Vasseur, F., Violle, C., Shipley, B., Enquist, B.J. and D. Vile. Testing models for the leaf economics spectrum with leaf and w …
39-
[2015 #183] Grady, J.M., Enquist,B.J., Dettweiler-Robinson, E., Wright, N. A., and F. A. Smith (2015) TECHNICAL COMMENTS: Response to Comments on “Evi …
40-
[2014 #191] Grady, J.M., Enquist, B.J., Dettweiler-Robinson, E., Wright, N.A. and F.A. Smith. (2014) Evidence for mesothermy in dinosaurs.Science, 344:1 …
41-
[2014 #192] Lamanna, C. A., Blonder, B., Violle, C. Kraft, N. J. B., Sandel, B. Simova, I. Donoghue, J., Svenning, J.C., McGill, B.J. Boyle, B. Dolins, …
42-
[2014 #194] Marquet,P.A., A.P. Allen, , J. H. Brown, J. Dunne, B.J. Enquist, J. Gillooly, P.A. Gowaty, J. L. Green, D. Storch, J. Harte, S. P. Hubbell, …
43-
[2014 #198] Blonder, B., Lamanna, C., Violle, C, and B.J. Enquist (2014) The n-dimensional hypervolume. Global Ecology and Biogeography, 23:595-609. …
44-
[2012 #213] Enquist, B. J., Boyle, B. (2012): SALVIAS – the SALVIAS vegetation inventory database. – In: Dengler, J., Oldeland, J., Jansen, F., Chytrý, …
45-
[2011 #225] Enquist, B.J. (2011) Forest annual carbon cost: comment. Ecology, 92:1994-1998. …
46-
[2011 #226] Kattge, J., Díaz, S., Lavorel S. et al. (2011) TRY – a global database of plant traits. Global Change Biology, 17:2905-2935. …
47-
[2011 #232] Stark, S.C., Bentley, L. P., and B.J. Enquist (2011) Response to Coomes & Allen (2009) ‘Testing the metabolic scaling theory of tree growth’ …
48-
[2009 #237] Price, C. A. and J. Enquist (2009) Comment on Coomes et al. “Scaling of xylem vessels and veins within the leaves of oak species”. Biology L …
49-
[2009 #240] Poore, B., Lamanna, C., Ebersole, J.J., and J. Enquist. (2009) Controls on radial growth of Mountain Big Sagebrush and implications for clim …
50-
[2007 #253] Enquist, B.J. (2007). Journal Club – An ecologist wonders how biotic feedback matters to global-change research. Nature 450:139. …
51-
[2007 #255] Enquist, B.J. and S.C. Stark (2007) Correspondence – Follow Thompson to make biology a capital-S Science. Nature 446:611. …
52-
[2003 #293] West, G.B., V. M. Savage, J. Gillooly, J. Enquist, William H. Woodruff and James H. Brown. (2003). Brief Communication – But Why DoesMetabol …
53-
[2002 #307] West, G.B., Savage, V.M., Gillooly, J., Enquist, B.J. Woodruff, W.H. and J.H. Brown (2002). Red herrings and rotten fish. arXiv:physics/0211 …
54-
[1999 #321] Enquist, B.J., Brown, J.H. & West, G.B. (1999) Plant energetics and population density Response. Nature 398:572-573. …
9+
Papers in public Google Doc but NOT found in HTML: 0
5510

5611
DOI FAILURES / ANOMALIES
5712
----------------------------------------

chat_provenance_log.md

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -327,3 +327,13 @@ Outcome: Completed coordinated tri-agent review-only audit (no code edits). Cros
327327
Prompt: Update research tab to say closed-loop and open-path chamber-based approaches; add co2fluxtent (https://github.com/PaulESantos/co2fluxtent) alongside fluxible in research.md flux tooling paragraph and as a new section in software.md.
328328
Source session: VS Code Copilot Chat
329329
Outcome: (1) research.md flux measurement sentence updated to cite both closed-loop and open-path approaches; (2) co2fluxtent described as a companion open-path package after the fluxible paragraph; (3) software.md new co2fluxtent section added below fluxible with GitHub link and install snippet.
330+
331+
43. Date: 2026-06-06
332+
Prompt: For https://enquistlab.github.io/publications/ add the missing 2021 Nature paper "How deregulation, drought and increasing fire impact Amazonian biodiversity" and ensure it groups under Tropical Ecology and Conservation Impacts.
333+
Source session: VS Code Copilot Chat
334+
Outcome: Added a targeted repair path in `scripts/sync_publications_html.py` for the missing Feng et al. (2021) Nature citation, reran the sync so `_includes/publications_full_from_doc.md` now contains the paper, and expanded the `conservation-impacts` topic matchers in `_pages/publications.html` with Amazonian biodiversity / deregulation / fire signals. The paper now exists in the master publication list and is classified for both Tropical Ecology and Conservation Impacts. Updated `assets/cv/publications_sync_report.txt` during validation.
335+
336+
44. Date: 2026-06-06
337+
Prompt: All 2021 papers appear to be missing from the publications page; keep the website publications synced to the public Google Doc.
338+
Source session: VS Code Copilot Chat
339+
Outcome: Added `scripts/rebuild_publications_include_from_doc.py` to rebuild `_includes/publications_full_from_doc.md` from the public Google Doc HTML export while regrouping entries by inferred citation year, which restored an explicit 2021 year block and the missing 2021 papers. Updated `.github/workflows/sync-google-doc-cv.yml` to run the rebuild script before the sync checker. Adjusted `scripts/sync_publications_html.py` so its missing-paper report now compares the site include against the public Google Doc HTML source of truth. Final validation report: `Papers in public Google Doc but NOT found in HTML: 0`.
Lines changed: 233 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,233 @@
1+
#!/usr/bin/env python3
2+
"""
3+
Rebuild the publications include from the current cleaned include plus the
4+
public Google Doc HTML export.
5+
6+
This script fixes two recurring drift problems:
7+
1. Items can sit under the wrong year block if the Google Doc year heading
8+
formatting changes.
9+
2. New items can appear in the public Google Doc but not in the website include.
10+
11+
The strategy is conservative:
12+
- Start from the existing cleaned website include.
13+
- Reassign each existing <li> to the year inferred from the citation text.
14+
- Parse the public Google Doc HTML export, clean the anchor URLs, and add any
15+
missing items by year.
16+
- Render a normalized include with one explicit year block per inferred year.
17+
"""
18+
19+
import re
20+
import urllib.parse
21+
import urllib.request
22+
from collections import OrderedDict
23+
from html import escape, unescape
24+
from html.parser import HTMLParser
25+
from pathlib import Path
26+
27+
28+
DOC_ID = "1OMolpDfY6c73qRgUFgYh0XbMS8EykUoy_hvrI8fnTn0"
29+
DOC_HTML_URL = f"https://docs.google.com/document/d/{DOC_ID}/export?format=html"
30+
PUBS_HTML = Path("_includes/publications_full_from_doc.md")
31+
SECTION_LABEL = "Peer Reviewed Journal Aricles"
32+
SECTION_END_LABELS = (
33+
"Other Publications",
34+
"Book Chapters",
35+
"Books",
36+
)
37+
38+
SECTION_RE = re.compile(r"<p>\s*([^<]+?)\s*</p><ol>(.*?)</ol>", re.IGNORECASE | re.DOTALL)
39+
LI_RE = re.compile(r"<li>.*?</li>", re.IGNORECASE | re.DOTALL)
40+
YEAR_RE = re.compile(r"\((19|20)\d{2}\)|\b(19|20)\d{2}\b")
41+
42+
43+
def clean_google_href(href):
44+
href = unescape(href or "").strip()
45+
parsed = urllib.parse.urlparse(href)
46+
if parsed.netloc.endswith("google.com") and parsed.path == "/url":
47+
q = urllib.parse.parse_qs(parsed.query).get("q")
48+
if q:
49+
return q[0]
50+
return href
51+
52+
53+
def strip_tags(html_text):
54+
text = re.sub(r"<[^>]+>", "", html_text)
55+
text = text.replace("\xa0", " ")
56+
return re.sub(r"\s+", " ", unescape(text)).strip()
57+
58+
59+
def infer_year(html_text):
60+
text = strip_tags(html_text)
61+
match = YEAR_RE.search(text)
62+
if not match:
63+
return None
64+
year_text = match.group(0).strip("()")
65+
return int(year_text)
66+
67+
68+
def title_key(html_text):
69+
text = strip_tags(html_text)
70+
match = re.search(r"\(\d{4}\)[.\s]+(.+)", text)
71+
if match:
72+
text = match.group(1)
73+
text = re.sub(r"[^a-z0-9]+", " ", text.lower())
74+
return re.sub(r"\s+", " ", text).strip()[:160]
75+
76+
77+
def normalize_li(html_text):
78+
html_text = html_text.replace("\xa0", " ")
79+
html_text = re.sub(r"\s+", " ", html_text)
80+
html_text = re.sub(r">\s+", ">", html_text)
81+
html_text = re.sub(r"\s+<", "<", html_text)
82+
html_text = re.sub(r"\s+([,.;:)])", r"\1", html_text)
83+
html_text = re.sub(r"([(" + "\[] )\s+", r"\1", html_text)
84+
return html_text.strip()
85+
86+
87+
class DocPublicationParser(HTMLParser):
88+
def __init__(self):
89+
super().__init__()
90+
self.items = []
91+
self.in_li = False
92+
self.parts = []
93+
94+
def handle_starttag(self, tag, attrs):
95+
if tag == "li":
96+
self.in_li = True
97+
self.parts = []
98+
return
99+
100+
if self.in_li and tag == "a":
101+
attr_map = dict(attrs)
102+
href = clean_google_href(attr_map.get("href", ""))
103+
self.parts.append(
104+
'<a href="{}" target="_blank" rel="noopener noreferrer">'.format(
105+
escape(href, quote=True)
106+
)
107+
)
108+
109+
def handle_endtag(self, tag):
110+
if self.in_li and tag == "a":
111+
self.parts.append("</a>")
112+
return
113+
114+
if tag == "li" and self.in_li:
115+
item = normalize_li("<li>" + "".join(self.parts) + "</li>")
116+
if strip_tags(item):
117+
self.items.append(item)
118+
self.in_li = False
119+
self.parts = []
120+
121+
def handle_data(self, data):
122+
if self.in_li:
123+
self.parts.append(escape(data, quote=False))
124+
125+
def handle_entityref(self, name):
126+
if self.in_li:
127+
self.parts.append(escape(unescape(f"&{name};"), quote=False))
128+
129+
def handle_charref(self, name):
130+
if self.in_li:
131+
self.parts.append(escape(unescape(f"&#{name};"), quote=False))
132+
133+
134+
def load_doc_html():
135+
req = urllib.request.Request(DOC_HTML_URL, headers={"User-Agent": "publications-include-sync/1.0"})
136+
with urllib.request.urlopen(req, timeout=60) as resp:
137+
return resp.read().decode("utf-8", errors="replace")
138+
139+
140+
def extract_doc_section(doc_html):
141+
start = doc_html.find(SECTION_LABEL)
142+
if start == -1:
143+
raise RuntimeError(f"Could not find section label: {SECTION_LABEL}")
144+
145+
end_candidates = [doc_html.find(label, start + len(SECTION_LABEL)) for label in SECTION_END_LABELS]
146+
end_candidates = [idx for idx in end_candidates if idx != -1]
147+
end = min(end_candidates) if end_candidates else len(doc_html)
148+
return doc_html[start:end]
149+
150+
151+
def parse_doc_items_by_year(doc_html):
152+
parser = DocPublicationParser()
153+
parser.feed(extract_doc_section(doc_html))
154+
155+
items_by_year = OrderedDict()
156+
keys_by_year = {}
157+
for item in parser.items:
158+
year = infer_year(item)
159+
if year is None:
160+
continue
161+
key = title_key(item)
162+
if year not in items_by_year:
163+
items_by_year[year] = []
164+
keys_by_year[year] = set()
165+
if key and key not in keys_by_year[year]:
166+
items_by_year[year].append(item)
167+
keys_by_year[year].add(key)
168+
return items_by_year
169+
170+
171+
def parse_existing_include(include_text):
172+
items_by_year = OrderedDict()
173+
keys_by_year = {}
174+
175+
for match in SECTION_RE.finditer(include_text):
176+
section_year = match.group(1).strip()
177+
body = match.group(2)
178+
for li_match in LI_RE.finditer(body):
179+
item = normalize_li(li_match.group(0))
180+
inferred = infer_year(item)
181+
if inferred is None:
182+
try:
183+
inferred = int(section_year)
184+
except ValueError:
185+
continue
186+
key = title_key(item)
187+
if inferred not in items_by_year:
188+
items_by_year[inferred] = []
189+
keys_by_year[inferred] = set()
190+
if key and key not in keys_by_year[inferred]:
191+
items_by_year[inferred].append(item)
192+
keys_by_year[inferred].add(key)
193+
return items_by_year, keys_by_year
194+
195+
196+
def render_include(items_by_year):
197+
ordered_years = sorted(items_by_year.keys(), reverse=True)
198+
lines = [
199+
"<!-- Auto-generated from shared Google Doc: peer-reviewed publication section -->",
200+
f"<p>{SECTION_LABEL}</p>",
201+
]
202+
for year in ordered_years:
203+
items = "".join(items_by_year[year])
204+
lines.append(f"<p>{year}</p><ol>{items}</ol>")
205+
return "\n".join(lines) + "\n"
206+
207+
208+
def main():
209+
include_text = PUBS_HTML.read_text(encoding="utf-8") if PUBS_HTML.exists() else ""
210+
existing_by_year, keys_by_year = parse_existing_include(include_text)
211+
212+
doc_html = load_doc_html()
213+
doc_by_year = parse_doc_items_by_year(doc_html)
214+
215+
additions = 0
216+
for year, items in doc_by_year.items():
217+
if year not in existing_by_year:
218+
existing_by_year[year] = []
219+
keys_by_year[year] = set()
220+
for item in items:
221+
key = title_key(item)
222+
if key and key not in keys_by_year[year]:
223+
existing_by_year[year].append(item)
224+
keys_by_year[year].add(key)
225+
additions += 1
226+
227+
rendered = render_include(existing_by_year)
228+
PUBS_HTML.write_text(rendered, encoding="utf-8")
229+
print(f"[include] Wrote {PUBS_HTML} with {len(existing_by_year)} year blocks and {additions} added item(s).")
230+
231+
232+
if __name__ == "__main__":
233+
main()

0 commit comments

Comments
 (0)