## 独立词库 理解词原先只能跟着课程单元走,学完 A0 十课词汇量只增加约 22 个实词, 不足以解决"记不住单词"。新增一份独立词库 assets/words/wordbank.json (2748 词,A1–B1),挂进 receptiveWordRegistry 的合成单元 bank-A1/A2/B1, 完全复用理解词已有的状态机,不依赖课程进度,第一天就能用。 数据来源、许可与合成规则记在 tool/words/DATA-NOTE.md:CEFR-J 定等级、 公开词书提供音标、AI 重写全部释义并生成例句、OpenSubtitles 提供口语词频。 词书部分为 CC BY-NC-SA 4.0 且上游权利不明,仅供个人非商用; 若要分发或上架,须替换音标那一列。 ## 背单词机制 - 间隔阶梯 1/3/7/15/30/60/120 天,连续答对上一级,答错回第一级。 原先首次答对后要等 7 天才复习,正是"第二天就忘"的成因。 - 每日新词上限(10 分钟 8 个 / 20 分钟 15 个 / 30 分钟 20 个)。 阶梯第一级是次日,今天引入的新词就是明天的工作量。 - 新词按口语频率发放,不再按字母序 —— A1 从 a.m./ability 变成 no/not/know/just。 - 三个方向按层级轮转:看词(英→中)→ 听词(音→中)→ 想词(中→英)。 想词题仍是选择题,不要求产出,理解词定位不变,不进升级分母。 - 单词页独立成 tab,首页今日任务卡下方给一张认词入口卡。 ## 用法对照 课程 JSON 增加 usage 字段(when/reply/swap/confuse):一个句型用在什么场合、 对方通常怎么答、还能怎么说、跟哪个学过的句型容易混。 知道 How are you? 的意思,不等于知道它不是用来问名字的。 ## 复习流 - 当日快闪(recap)独立成队列,不占复习预算,也不计入积压。 - 只发放当日预算内的量,其余保持到期状态等下次,不悄悄丢弃或改期。 - 答错的项隔几题后回来,而不是立刻重问。 测试 296 通过,flutter analyze 干净。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
70 lines
2.1 KiB
Python
70 lines
2.1 KiB
Python
"""Writes a frequency rank onto every bank word.
|
|
|
|
Without this the bank is handed out in alphabetical order, so the first two
|
|
weeks are nothing but words starting with `a`. Rank 0 is the most common word.
|
|
|
|
A multi-word entry (`bus stop`, `air conditioning`) is not in a word frequency
|
|
list at all, so it takes the rank of its rarest part: a phrase is no easier
|
|
than the hardest word in it.
|
|
"""
|
|
|
|
import json
|
|
import os
|
|
import re
|
|
import sys
|
|
|
|
ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
|
BANK = os.path.join(ROOT, 'assets/words/wordbank.json')
|
|
|
|
# Words the list does not have at all go last, but stay in their level.
|
|
UNRANKED = 99999
|
|
|
|
|
|
def load_ranks(path):
|
|
ranks = {}
|
|
with open(path) as handle:
|
|
for position, line in enumerate(handle):
|
|
parts = line.split()
|
|
if len(parts) == 2:
|
|
ranks.setdefault(parts[0], position)
|
|
return ranks
|
|
|
|
|
|
def rank_of(word, ranks):
|
|
direct = ranks.get(word.lower())
|
|
if direct is not None:
|
|
return direct
|
|
parts = [re.sub(r"[^a-z']", '', part) for part in word.lower().split()]
|
|
parts = [part for part in parts if part]
|
|
found = [ranks[part] for part in parts if part in ranks]
|
|
if len(found) == len(parts) and found:
|
|
return max(found)
|
|
return UNRANKED
|
|
|
|
|
|
def main():
|
|
frequency = sys.argv[1]
|
|
ranks = load_ranks(frequency)
|
|
with open(BANK) as handle:
|
|
bank = json.load(handle)
|
|
unranked = 0
|
|
for word in bank['words']:
|
|
rank = rank_of(word['en'], ranks)
|
|
word['rank'] = rank
|
|
if rank == UNRANKED:
|
|
unranked += 1
|
|
with open(BANK, 'w') as handle:
|
|
json.dump(bank, handle, ensure_ascii=False, separators=(',', ':'))
|
|
handle.write('\n')
|
|
print('ranked %d, unranked %d' % (len(bank['words']) - unranked, unranked))
|
|
for level in ('A1', 'A2', 'B1'):
|
|
words = sorted(
|
|
(w for w in bank['words'] if w['level'] == level),
|
|
key=lambda w: w['rank'],
|
|
)
|
|
print('%s first 12: %s' % (level, ' '.join(w['en'] for w in words[:12])))
|
|
|
|
|
|
if __name__ == '__main__':
|
|
main()
|