## 独立词库 理解词原先只能跟着课程单元走,学完 A0 十课词汇量只增加约 22 个实词, 不足以解决"记不住单词"。新增一份独立词库 assets/words/wordbank.json (2748 词,A1–B1),挂进 receptiveWordRegistry 的合成单元 bank-A1/A2/B1, 完全复用理解词已有的状态机,不依赖课程进度,第一天就能用。 数据来源、许可与合成规则记在 tool/words/DATA-NOTE.md:CEFR-J 定等级、 公开词书提供音标、AI 重写全部释义并生成例句、OpenSubtitles 提供口语词频。 词书部分为 CC BY-NC-SA 4.0 且上游权利不明,仅供个人非商用; 若要分发或上架,须替换音标那一列。 ## 背单词机制 - 间隔阶梯 1/3/7/15/30/60/120 天,连续答对上一级,答错回第一级。 原先首次答对后要等 7 天才复习,正是"第二天就忘"的成因。 - 每日新词上限(10 分钟 8 个 / 20 分钟 15 个 / 30 分钟 20 个)。 阶梯第一级是次日,今天引入的新词就是明天的工作量。 - 新词按口语频率发放,不再按字母序 —— A1 从 a.m./ability 变成 no/not/know/just。 - 三个方向按层级轮转:看词(英→中)→ 听词(音→中)→ 想词(中→英)。 想词题仍是选择题,不要求产出,理解词定位不变,不进升级分母。 - 单词页独立成 tab,首页今日任务卡下方给一张认词入口卡。 ## 用法对照 课程 JSON 增加 usage 字段(when/reply/swap/confuse):一个句型用在什么场合、 对方通常怎么答、还能怎么说、跟哪个学过的句型容易混。 知道 How are you? 的意思,不等于知道它不是用来问名字的。 ## 复习流 - 当日快闪(recap)独立成队列,不占复习预算,也不计入积压。 - 只发放当日预算内的量,其余保持到期状态等下次,不悄悄丢弃或改期。 - 答错的项隔几题后回来,而不是立刻重问。 测试 296 通过,flutter analyze 干净。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
94 lines
3.0 KiB
Python
94 lines
3.0 KiB
Python
"""Merges the AI rewrite back into assets/words/wordbank.json.
|
||
|
||
Anything the model returned is checked before it lands: a gloss that is empty,
|
||
parenthesised or sentence-long is worse than the word book's version, so the
|
||
old value is kept and reported instead of being written over.
|
||
"""
|
||
|
||
import json
|
||
import os
|
||
import re
|
||
import sys
|
||
import tempfile
|
||
|
||
ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||
BANK = os.path.join(ROOT, 'assets/words/wordbank.json')
|
||
|
||
BAD_GLOSS = re.compile(r'[()()\[\]【】<>]')
|
||
|
||
|
||
def usable_gloss(text):
|
||
text = (text or '').strip()
|
||
return bool(text) and len(text) <= 10 and not BAD_GLOSS.search(text)
|
||
|
||
|
||
def mentions(word, sentence):
|
||
"""Whether the sentence visibly uses the word.
|
||
|
||
This cannot be a hard check: `teach` shows up as "taught" and `bus stop`
|
||
as two words, so a prefix match has no way to confirm them. It only flags
|
||
a sentence as worth a human glance -- nothing is dropped over it.
|
||
"""
|
||
parts = [re.sub(r'[^a-z]', '', part) for part in word.lower().split()]
|
||
stem = max(parts, key=len)[:3]
|
||
return bool(stem) and stem in sentence.lower()
|
||
|
||
|
||
def main():
|
||
cache_path = os.environ.get(
|
||
'ENRICH_CACHE', os.path.join(tempfile.gettempdir(), 'enrich_cache.jsonl')
|
||
)
|
||
rows = {}
|
||
with open(cache_path) as handle:
|
||
for line in handle:
|
||
row = json.loads(line)
|
||
rows[row['id']] = row
|
||
with open(BANK) as handle:
|
||
bank = json.load(handle)
|
||
|
||
kept = {'gloss': 0, 'example': 0, 'missing': 0, 'bad_gloss': 0, 'no_example': 0}
|
||
unclear = []
|
||
for word in bank['words']:
|
||
row = rows.get(word['id'])
|
||
if row is None:
|
||
kept['missing'] += 1
|
||
continue
|
||
if usable_gloss(row.get('zh')):
|
||
word['zh'] = row['zh'].strip()
|
||
more = (row.get('more') or '').strip()
|
||
senses = [s for s in more.split(';') if usable_gloss(s) and s != word['zh']]
|
||
if senses:
|
||
word['more'] = ';'.join(senses[:2])
|
||
else:
|
||
word.pop('more', None)
|
||
kept['gloss'] += 1
|
||
else:
|
||
kept['bad_gloss'] += 1
|
||
sentence = (row.get('ex') or '').strip()
|
||
if not sentence:
|
||
kept['no_example'] += 1
|
||
continue
|
||
word['ex'] = sentence
|
||
word['exZh'] = (row.get('ex_zh') or '').strip()
|
||
kept['example'] += 1
|
||
if not mentions(word['en'], sentence):
|
||
unclear.append('%s | %s' % (word['en'], sentence))
|
||
|
||
print(json.dumps(kept, indent=2))
|
||
if unclear:
|
||
print('%d sentences to eyeball:' % len(unclear))
|
||
for line in unclear:
|
||
print(' ' + line)
|
||
if '--dry-run' in sys.argv:
|
||
return
|
||
bank['source'] = (
|
||
'CEFR-J Vocabulary Profile 1.5 (levels) + 公开词书 (音标) + AI 重写 (释义与例句)'
|
||
)
|
||
with open(BANK, 'w') as handle:
|
||
json.dump(bank, handle, ensure_ascii=False, separators=(',', ':'))
|
||
handle.write('\n')
|
||
|
||
|
||
if __name__ == '__main__':
|
||
main()
|