英語コーパス研究
Online ISSN : 2759-5676
Print ISSN : 1340-301X
最新号
選択された号の論文の10件中1~10を表示しています
論文
  • Yuichiro KOBAYASHI
    2026 年33 巻 p. 1-16
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    With the significant shift in the second language (L2) education landscape, developing learners’ productive skills, particularly speaking and writing, is increasingly emphasized. Consequently, there is a growing demand for efficient and objective assessment methods. To enhance the transparency and interpretability of automated language assessment, this study establishes a framework that leverages key feature analysis with the Boruta algorithm. The study aims to provide clearer insights into the linguistic features driving assessment decisions while maintaining high predictive accuracy. To achieve this, data from the Longitudinal Corpus of L2 Spoken English are used, specifically, results from the Telephone Standard Speaking Test (TSST), which is a phone-based speaking test. Boruta is applied to identify key linguistic features predictive of L2 speaking proficiency, systematically distinguishing genuine predictors from statistical noise. Focusing on TSST Levels 3-6 (CEFR A2-B1; n = 821), Boruta’s effectiveness in identifying a concise and meaningful feature set is demonstrated. The model achieves a 71.36% mean accuracy and 0.651 correlation (p < .001) with human ratings, thus revealing distinct patterns in linguistic feature use across proficiency levels. Features such as vocabulary, grammar, syntax, and discourse exhibit particularly strong relationships with proficiency. Although high-performing, traditional random forest can pose challenges for objective interpretation of variable importance, Boruta allows for statistical validation of the selected features. This study contributes to a deeper understanding of L2 development and assessment by balancing accuracy and interpretability, which enhances objectivity and consistency, thereby providing pedagogical insights that are not accessible through more complex and less transparent methods.

  • Tatsuya ISHII, Yohei HIRANO
    2026 年33 巻 p. 17-34
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    This study investigates whether large language models (LLMs) can expedite reliable move annotation-where “moves” are functional discourse units performing distinct communicative purposes-for spoken academic discourse, with the goal of deriving a pedagogically useful, move-based phraseological list from Three-Minute Thesis (3MT) presentations. A corpus of 160 finalist/winner 3MT talks from 12 universities (69,705 tokens; 7,846 types) was transcribed from YouTube captions and segmented into eight moves (Orientation, Rationale, Framework, Purpose, Methods, Results, Implication, and Termination) using ChatGPT-o1 with a fixed prompt. Two experienced EAP raters independently verified sentence-level labels; residual disagreements were adjudicated by two additional instructors. Agreement with human labels was near-expert (Cohen’s κ = 0.953 vs. Rater 1; 0.924 vs. Rater 2), matching or exceeding human-human alignment (κ = 0.905); three-coder reliability was strong (Krippendorff’s α = 0.927). Only 1.4% of sentences (55/3,914) required reconsideration, and accuracy remained ≥ 0.968 across moves, with the main difficulty at Purpose boundaries. Using the validated move corpus, we extracted four-word phrase-frames (one open slot) per move via text dispersion keyness (log-likelihood based on text frequency) with a file-frequency cutoff (≥ 3 talks). The resulting high-dispersion phrase-frames are strongly related to the functions of moves-for example, Orientation launchers (e.g., my name is *, I’m going to *), Purpose aim frames (e.g., my research * to), Methods chaining (e.g., to do this *, and then we *), Results report shells (e.g., we found that *), and Implication outlooks (e.g., in the future *), while generic function-word skeletons are de-emphasized. These findings demonstrate that an LLM-assisted, human-verified workflow can efficiently produce a reliable move corpus for 3MT and yield pedagogically oriented, text-dispersion-based phraseological resources. Limitations and avenues for refinement of prompt design, broader sentence-level improvements, and application to other genres are discussed.

  • Hiromitsu FUKUMOTO
    2026 年33 巻 p. 35-58
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    This article investigates the diachronic behavior, stylistic distribution, and socio-historical patterning of split infinitives in the U.S. State of the Union (SOTU) addresses. Using a corpus of SOTU addresses spanning 1790-2024 (Washington to Biden), compiled from The American Presidency Project and annotated with Brill’s part-of-speech tagger, 220 split infinitive tokens attested in 1846-2024 are identified; the earliest occurs in Polk’s 1846 address. These tokens are analyzed with respect to three questions: (i) changes in frequency and distribution; (ii) stylistic and rhetorical differences across presidents and eras; and (iii) the lexical composition of split infinitives and its relation to broader socio-historical developments. The results show a non-linear trajectory: split infinitives are absent from early SOTU texts, first appear in Polk, cluster under late-nineteenth-century presidents such as Grant and Cleveland, decline sharply in the prescriptively dominated early twentieth century, and re-emerge and stabilize from the post-war period onward. Split infinitives are unevenly distributed across presidents and eras, forming part of stylistic profiles. These profiles include earnest recommendations in the nineteenth century, “proper” and “constructive” administration in the Progressive Era, managerial and strategic framing in the Cold War period, and completion-focused and stance-oriented narratives in contemporary addresses. At the lexical level, the corpus exhibits diversity (104 adverbs, 158 verbs) but concentration around templates such as to further aid, to better secure, to fully fund, and to finally end, which recur in shifting policy domains from Reconstruction-era civil rights to homeland security and space exploration. The syllable-level annotation suggests that many of these templates exploit prosodic patterns that favor spoken delivery. This study extends corpus-based work on the split infinitive by providing a detailed, single-genre history and by showing how split infinitive templates function as discourse strategies in the American presidency.

  • 西垣 知佳子, 川名 隆行, 中井 康平, 見目 慎也, 山崎 達也
    2026 年33 巻 p. 59-71
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    This exploratory study investigated which learner factors are associated with junior high school students’ intention to use a data-driven learning (hDDL) application beyond class time after initial classroom implementations. Participants were 112 first-year students at a Japanese university-affiliated lower secondary school. In February 2025, students experienced four hDDL-supported lessons using a controlled example-sentence corpus designed for CEFR A1-A2 learners. Immediately after the implementation, students completed a 19-item Likert-scale questionnaire targeting ease of use, perceived usefulness across skills, classroom learning preferences, and English-skill preferences. Using JASP, we first reported descriptive statistics for all items and examined internal consistency for theoretically motivated composites. A usefulness composite (mean of Q3-Q9; α=.851) and an English-preference composite (mean of Q13-Q17; α=.904) were retained, whereas low-reliability sets were analyzed at the item level. Associations between the behavioral intention item (Q2; intention to use hDDL outside class) and other variables were examined using Spearman’s rank correlations with false discovery rate control. To identify independent predictors of intention, we conducted proportional-odds ordinal logistic regression with standardized predictors and reported odds ratios with 95% confidence intervals. Results showed that intention to use hDDL outside class was most strongly related to perceived usefulness and was independently predicted by usefulness (OR=3.48, 95% CI [2.07, 5.83]) and English preference (OR=2.03, 95% CI [1.20, 3.44]). Preference for learning through peer discussion exhibited a small negative effect after controlling other factors (OR=0.66, 95% CI [0.45, 0.97]). These findings suggest that, for entry-level learners, enhancing concrete perceptions of usefulness across language skills is central to fostering sustained engagement with DDL tools, while instructional designs may need flexible participation structures to accommodate differing classroom interaction preferences. The study provides an initial learner-profile perspective to inform scalable DDL integration in lower secondary EFL contexts.

  • Yuki SUGAWARA
    2026 年33 巻 p. 73-89
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    Understanding how the verb explain functions across natural language is crucial for philosophy of explanation, corpus linguistics, and science and technology studies, yet existing approaches have lacked a comprehensive, distributional account of explanatory practices at scale. This study proposes Distributional Concept Analysis (DCA), a multilayered framework that integrates E5 sentence embeddings, UMAP-based semantic galaxy visualization, PCA-derived semantic axes, the Combined Topic Model (CTM), and GPT-5-assisted topic labeling. Using 10,000 instances of explain drawn from the British National Corpus, DCA constructs a high-dimensional “semantic galaxy” in which explanatory uses form distinct but interconnected star-like clusters corresponding to scientific, institutional, interactional, and narrative contexts. PCA reveals two continuous semantic axes-one ranging between scientific and narrative, and the other between institutional and interpersonal-that organize this galaxy globally, providing a distributed structure that extends beyond genre boundaries. CTM and GPT-5 further identify six coherent explanatory constellations, each defined by distinctive lexical, grammatical, and discourse features, which together demonstrate that explain serves as a hub linking diverse explanatory modes. These findings show that explanation is not a monolithic linguistic or conceptual category but a multi-centered semantic ecology shaped by genre, discourse aims, and interactional conditions. Conceptually, the study extends Moretti’s notion of distant reading by introducing cosmic reading, an approach that observes not only textual patterns but the semantic landscapes and constellations emergent from large-scale embeddings. By integrating quantitative semantics with generative interpretation, DCA operationalizes Machery’s call for empirical philosophy of science and broadens Overton’s corpus-based study of scientific explanation to encompass the full range of everyday explanatory practices. This work thus provides a new methodological foundation for analyzing explanation as a distributed, ecologically structured phenomenon in natural language.

  • Shusaku NAKAYAMA
    2026 年33 巻 p. 91-107
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    It is worthwhile for learners to master collocations, two or more words that often go together, because they can help enhance speaking fluency and save one’s memory resources. Their high pedagogical value has inspired researchers to create collocation lists and dictionaries. However, most of them are geared toward academic contexts and generally have an overwhelming number of collocations, making it difficult to determine which collocations learners, especially novices, should study first. This research thus seeks to develop a list of basic collocations for spoken contexts, intended for learners with limited collocational knowledge. To this end, collocations were compiled that consisted solely of words in the New General Service List-Spoken 1.2, a list of the most frequent 721 words in general spoken English, providing up to 90% coverage of general spoken texts. Collocations were extracted using three statistical parameters: frequency of occurrence, strength of word association, and directionality, a parameter for judging which of a collocation’s components is the node word. The Online OXFORD Collocation Dictionary of English was consulted to ensure that the collocation list contained only those whose components were semantically and lexically connected; any that did not appear in this publication were removed. The completed list comprised 1,138 different collocations. Results showed that frequently-used collocations were consistently those with prevalent word class patterns, while strongly connected components were not necessarily prevalent. Analyzing the collocations quantitatively revealed that verb-preposition and noun-preposition collocations were more prevalent across spoken than written contexts, while verb-noun and adjective-noun collocations were prevalent across all contexts. The list is smaller and more manageable than existing collocation lists, possibly making it a suitable resource for novice learners with only a basic spoken vocabulary. Future research should investigate the validity of the collocation list and how to teach collocations effectively to learners.

  • 藤本 和子
    2026 年33 巻 p. 109-126
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    Previous corpus-based studies have reported that the present perfect is more frequently used without adverbials. In contrast, in pedagogical contexts, it is often taught in association with adverbials. This study aims to provide teaching insights for Japanese university students on the co-occurrence patterns of the present perfect with adverbials. The analysis focuses on such patterns in paragraph writing by Japanese university students at CEFR A2 and B1 levels, as well as in example sentences from university-level ELT grammar materials and high school English textbooks. The results reveal that university students used the present perfect with adverbials more frequently in their writing. Overall, ELT materials provide more examples of the present perfect co-occurring with adverbials, and all the high school textbooks examined show a similarly high frequency of such co-occurrence. While pedagogical grammar can help learners’ understanding, instruction that overemphasizes the co-occurrence relationships may prevent learners from using the present perfect independently of adverbials. At the university level, it may be necessary to supplement explanations in teaching materials by raising students’ awareness that the present perfect is more often used without adverbials and by providing sufficient contextualized examples to ensure a thorough understanding of its meanings.

  • Yuanyuan XIAO
    2026 年33 巻 p. 127-146
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    This study investigates subjective bias in English political news by combining corpus linguistics and critical discourse analysis methodologies. While news reporting is expected to be neutral, it is often influenced by journalists’ perspectives and word choices, resulting in intentional or unconscious biases. Focusing on English political news, this research systematically analyzes how linguistic choices reflect biases and affect readers’ perceptions. This study proceeds in three stages: (1) identifying topics in English political news using BERTopic, a data-driven topic modeling tool that minimizes researcher subjectivity, (2) uncovering deeper linguistic patterns and bias markers through a BERT-based bias neutralization model (Pryzant et al., 2020), and (3) applying Fairclough’s three-dimensional framework to analyze the sociocultural, political, and ideological forces shaping media discourse. This integrated methodology allows for both quantitative precision and contextual interpretation.

    Findings reveal distinct patterns in how different news outlets frame similar topics, with national ideologies and editorial stances influencing linguistic choices and citation strategies. Analysis of news outlets like The Jerusalem Post and Tehran Times highlights significant biases driven by national ideologies, while comparisons between Express and The Guardian reveal different editorial stances. Further analysis of The New York Times and The Washington Post uncovers differences in language and framing despite similar levels of bias. These results demonstrate how news media outlets construct narratives aligned with their ideological and institutional contexts, emphasizing the importance of language in shaping public perceptions.

研究ノート
  • 安間 一雄
    2026 年33 巻 p. 147-159
    発行日: 2026年
    公開日: 2026/06/25
    ジャーナル オープンアクセス

    Managing text types is an integral part in teaching/learning reading comprehension and essay writing. Without a good understanding and implementation of strategies for text types success in academic production as well as business handling would be hopeless. EFL textbooks ─ here we focus on MEXT-approved senior high school textbooks ─ do not appear to take into consideration a good balance of texts of various types. In this research 10 textbooks of relatively high market share were examined as to how varied their texts are in terms of four text types: expository, descriptive, persuasive, and narrative. The result was an immense degree of difference in text type orientations. One of our findings was the major contrast in expository versus narrative, which is often the focus of discussion in second language acquisition. With respect to readiness for university-level academic English education cultivating CALP as its basic element, selection of a textbook with a balanced repertoire in text types sounds rational and fruitful.

ソフトウェア紹介
feedback
Top