[{"content":"Consider two hypothetical sites.\nSite A is under a “Caution” algal bloom alert today, but its seven-day forecast line points downward. Site B has no alert yet, but its forecast line points upward. Which should we trust more?\nThe answer is not to choose one. An algal bloom alert describes the observed present, while an AI or numerical-model forecast describes a future calculated by a model. They answer questions about different points in time, so a difference between them does not mean that one is wrong.\nFigure 1. Algal alerts and AI forecasts answer questions about different times.\nThe 30-Second Takeaway # An algal bloom alert reports the current state from observations; AI and numerical-model forecasts estimate how it may change. In 2026, algal bloom forecasts provide information for the next seven days at 13 drinking-water source sites, but alerts and forecasts may concern different targets and times. To read algal bloom information well, check the forecast target, site, reference time, horizon, and uncertainty together. 1. “Caution” and “Warning” Turn Today’s Observations into Alert Levels # An algal bloom alert level is not assigned directly by a forecasting algorithm. It is issued by comparing samples collected and analyzed at monitoring sites with thresholds defined by law.\nUnder the criteria for drinking-water source zones in the Enforcement Decree of the Water Environment Conservation Act in force in March 2026, “Caution” (관심) applies when the cyanobacterial cell count is at least 1,000 cells/mL in two consecutive observations. “Warning” (경계) applies when the cyanobacterial cell count is at least 10,000 cells/mL in two consecutive observations or algal toxin is at least 10 μg/L. “Major Bloom” (조류대발생) applies when the cyanobacterial cell count is at least 1,000,000 cells/mL in two consecutive observations.[2]\nThree points are easy to miss in this structure.\nAn alert begins with an observed value, not a forecast. A single value above a threshold is not the same as an official alert. The criteria for issuing and lifting alerts look at two consecutive observations. Even when values are all described as measures of an “algal bloom,” cyanobacterial cell counts, algal toxin, and chlorophyll-a concentration are different measurement targets. Do not see the label “Caution” and read it as a model predicting that tomorrow will also be at the caution level. First check what was measured, where, and when to produce that label.\nFor example, suppose the algal-toxin threshold is not exceeded and the current cyanobacterial cell count is 12,000 cells/mL, but the previous observation was below 10,000 cells/mL. The requirement of two consecutive observations for “Warning” has not yet been met. That is why a recent measurement and the official alert level can appear inconsistent. Understanding this timing rule before reading the color coding prevents a misreading of the number.\n2. AI and Numerical Models Ask About the Next Seven Days # In May 2026, Korea’s National Institute of Environmental Research began using AI alongside existing methods for algal bloom forecasting. The plan was to combine an existing three-dimensional numerical model with an AI model trained on historical water-quality, water-quantity, and meteorological data, and to provide information about algal bloom occurrence over the following seven days. The number of drinking-water source sites covered increased from 9 to 13, and forecast information is published on the Mulmoa platform every Monday and Thursday from May through October.[1]\nUsing numerical models and AI together is a natural structure. A numerical model calculates water flows and physical dynamics; AI learns recurring patterns from historical data. Combining them does not make the future certain. The observations and weather forecasts used as inputs, the model structures, and the periods and sites included in training all leave their imprint on the result.\nIf the two models point in different directions, do not average them into one number and hide the disagreement. Treat it as a signal to investigate which differences in inputs and assumptions produced the divergence. Disagreement between models is inconvenient information, but it is not noise that should be concealed. If we record the conditions under which each model was wrong when the next observation arrives, forecasting becomes a validation record rather than a one-off attempt to guess the right answer.\nFigure 2. An alert reports the observed present; a forecast shows how a model expects conditions to change.\nDistinguishing the forecast target is important here. In a 2025 study published in the Journal of Cleaner Production, H. Jeon and I used weekly time series from multiple freshwater monitoring sites between 2016 and 2022 to predict chlorophyll-a concentration with a one-dimensional convolutional neural network. We also repeated CAM 1,000 times across different initializations so that a single explanation would not depend too heavily on chance.[3]\nThe chlorophyll-a concentration predicted by this model, however, is not the same target as the cyanobacterial cell count or algal toxin used in the current alert system. Reading the study as a direct prediction of alert levels would change the question the model answered.\n3. When Alerts and Forecasts Differ, Read Them as a Combination # Separating the algal bloom alert into current state and the seven-day forecast into future direction produces four combinations. The following table is a simplified exercise in judgment; it does not replace official alert actions.\nCurrent observation or alert Seven-day forecast direction How to read the combination Low Rising The absence of an alert now does not conflict with the possibility of a future increase. Treat it as an early signal and check it against the next observation. High Falling A forecast decline does not cancel the current alert. Record both the observed present and the expected direction of improvement. High Rising Current observations and the forecast direction send the same signal. You still need to confirm the site and target to which each applies. Low Falling The information is relatively consistent for that site and forecast horizon. It does not mean that other sites or times beyond the forecast horizon are safe. It is not enough to look only at whether the forecast line rises or falls. We also need to see how far it is from the alert threshold, when within the forecast horizon its direction changes, and when the latest observation was incorporated. A number close to a threshold may be colored “not Caution” or “Caution,” but that does not make the two states separate worlds.\n4. Four Things to Check Before Trusting an “Algal Bloom Forecast” # First, check what was predicted. Cyanobacterial cell count, algal toxin, chlorophyll-a concentration, and alert level are not interchangeable. The headline “algal bloom forecast” hides this difference.\nSecond, check which place and time the information describes. Even within the same river or lake, values can differ across monitoring sites. Distinguishing the observation date, publication date, and forecast reference time makes the word “latest” concrete.\nThird, check how far ahead the forecast extends. An observation today, a forecast for tomorrow, and a forecast for day seven do not carry the same weight. It is better to read the direction for each date and the update cycle together than to look only at the final point in the forecast horizon.\nFourth, ask whether uncertainty is visible. If only one number is provided, look for whatever information is available among a prediction interval, past error, differences between AI and numerical models, and distance from the alert threshold. If some of these are not shown, do not invent values to fill the gaps. Leave them as limitations that cannot be checked.\nFigure 3. Before trusting the label “algal-bloom forecast,” ask what, where, when, and how certain.\nThese criteria do not apply only to forecasts from public agencies. Models that I have researched or provide through related services are validated only within the ranges defined by their training periods, monitoring sites, and prediction targets. Moving to a new site, climate condition, or observation method requires revalidation. My models are no exception.\nConclusion # There is nothing strange about a day when an algal bloom alert and an AI forecast appear to differ. One describes the state observed now; the other describes a future possibility.\nThe next time you see the phrase “algal bloom forecast,” check four things before the number: what was predicted, where and when, how far ahead, and with how much certainty. These questions are the starting point for reading an alert and an AI forecast together.\nRelated Articles # Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals Does AI Really Say ‘I Don’t Know’ When It Encounters Unfamiliar Data? — Dataset Shift and OOD Sources # [1] Ministry of Climate, Energy and Environment \u0026amp; National Institute of Environmental Research. (2026.05.04). Precision Forecasting of Summer Algal Blooms with Artificial Intelligence (AI). https://www.mcee.go.kr/home/web/board/read.do?boardId=1861480\u0026boardMasterId=939\u0026menuId=10598\n[2] Korean Law Information Center. Enforcement Decree of the Water Environment Conservation Act, Article 28 and Attached Table 3 (effective 2026.03.24). https://www.law.go.kr/LSW/lsInfoP.do?lsiSeq=284747\n[3] Lee, D., \u0026amp; Jeon, H. (2025). Reinforced explainable AI for algal bloom forecasting under climate change: A multi-run class activation mapping (CAM) approach. Journal of Cleaner Production, 529, 146805. https://doi.org/10.1016/j.jclepro.2025.146805\nThe author provides related services in this field through AI Korea. This article does not promote any particular product or service.\nPredictions always contain uncertainty. They must not be used as the sole basis for decisions on disease-control or environmental policy.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"August 29, 2026","externalUrl":null,"permalink":"/en/posts/2026-08-29-ai-algal-bloom-forecast-literacy/","section":"Writing","summary":"Algal bloom alerts describe the observed present; AI and numerical models estimate change over the next seven days. This article explains how to read the forecast target, site, reference time, horizon, and uncertainty together.","title":"AI Has Arrived in Algal Bloom Forecasting — How Should We Read ‘Caution,’ ‘Warning,’ and a Seven-Day Forecast?","type":"posts"},{"content":"Professor at Hankuk University of Foreign Studies · PhD in Engineering, KAIST · CEO of AI Korea Inc.\nRead the latest About ","date":"August 29, 2026","externalUrl":null,"permalink":"/en/","section":"Donghyun Lee’s Trustworthy AI Notes","summary":"Professor at Hankuk University of Foreign Studies · PhD in Engineering, KAIST · CEO of AI Korea Inc.\nRead the latest About ","title":"Donghyun Lee’s Trustworthy AI Notes","type":"page"},{"content":"Essays on trustworthy AI and infectious-disease and environmental forecasting.\n","date":"August 29, 2026","externalUrl":null,"permalink":"/en/posts/","section":"Writing","summary":"Essays on trustworthy AI and infectious-disease and environmental forecasting.\n","title":"Writing","type":"posts"},{"content":"Python Fundamentals for the AI Era #1\nLet me begin with a piece of code. Ask AI to “read numbers from a file and calculate their average,” and you may receive something as neat-looking as this:\nvalues = [] for line in open(\u0026#34;data.csv\u0026#34;): try: values.append(float(line)) except: pass print(sum(values) / len(values)) Run it, and it finishes without an error. It prints a plausible-looking number as the average.\nWhat makes this code dangerous? I will reveal the answer later in the article. For now, one hint: this code will never stop when it is wrong.\nThe 30-Second Takeaway # Code that finishes without an error is not verified code. It is code that does not yet appear to be wrong. Silent errors in AI-written code recur in three places: assumptions about data types, range boundaries, and swallowed exceptions—the places where AI cannot see your data or your intent. Before trusting the output, spend 30 seconds checking a small input by hand, counting how many values entered the calculation, and confirming whether the first and last values were included. Figure 1. Code that runs is not necessarily code that is correct.\n1. Silent Code Is More Dangerous Than Code That Raises an Error # In the previous article, I described two kinds of errors. A loud error that stops execution announces its own existence, so feeding the error message back to AI will resolve it in most cases. The problem is the silent error: code that runs to completion without an error and even produces a plausible value.\nThe previous article reached the conclusion that we therefore need the ability to verify. This article takes up the next question: Where should we look?\nFortunately, silent errors do not arise everywhere. AI writes syntactically correct code. But there are three things it cannot see: the actual data inside my file, the precise intent in my head, and the failure conditions that will arise in practice. Silent errors are born regularly in these three places. Because their locations are predictable, we can define verification points in advance.\nWhile grading in the classroom, I see this difference repeatedly. A student whose code raises an error at least knows with certainty that something is wrong. A student whose code runs usually feels sure that it is correct. That is why working but incorrect code is much harder to discover. There is no signal that it is wrong, and no one suspects it.\nFigure 2. Silent errors live in three places the AI cannot see.\n2. The First Hiding Place — Why the Largest Value Becomes “950” # Suppose we read prices from a CSV file and ask AI to find the highest one. It produces this code:\nprices = [\u0026#34;1200\u0026#34;, \u0026#34;950\u0026#34;, \u0026#34;15000\u0026#34;, \u0026#34;800\u0026#34;] print(max(prices)) The output is 950, even though 15000 is plainly present.\nValues read from the file are strings enclosed in quotation marks, and strings are ordered lexicographically rather than numerically. Just as a word beginning with “9” would come after one beginning with “1” in this ordering, \u0026quot;950\u0026quot; is judged greater than \u0026quot;15000\u0026quot;. The calculation succeeds; only the answer is wrong.\nThe reason AI writes code like this is simple: it cannot open and inspect my file. It has to assume whether the values are numbers or strings. Even when that assumption is wrong, Python helpfully continues the calculation. With four values, we notice the error by eye. If four hours of data contain 40,000 rows, no one may notice.\nThe check takes one line: use type() to inspect one value before calculating, or verify that the code converts values with float().\n3. The Second Hiding Place — A Seven-Day Average Calculated from Six Days # This code calculates the average temperature over the last seven days.\ntemps = [27, 29, 31, 30, 28, 26, 25] week = temps[0:6] print(sum(week) / 7) The output is 24.42..., which looks plausible. The correct answer is 28.0.\nThe Python slice [0:6] returns six values because the ending index is excluded. The final day, 25 degrees, silently disappears, but the sum is still divided by 7, distorting the average twice over. The code has no reason to stop.\nThis kind of boundary mismatch occurs more often not in the first draft written by AI, but when a person edits AI-generated code. When requirements change to “through yesterday” or “excluding the first week,” a boundary can easily shift by one, and a shifted boundary does not raise an error. Neither AI nor Python can inspect the intent in our head.\nAgain, the check takes one line: print len(week) to count the values, then confirm that the first and last values are the ones you intended.\n4. The Third Hiding Place — A try Block That Swallows Errors (the Answer to the Opening Code) # Now return to the opening code. The danger lies in these two lines:\nexcept: pass They mean: “If something goes wrong, just move on.” Whether the file contains a header, a blank line, or an entry marked “N/A,” the code silently discards that line and calculates the average from whatever remains. Even if 400 of 1,000 lines are discarded, the only output is one perfectly ordinary-looking number. No one—not even the code itself—knows how many lines were lost.\nIn my experience, this pattern is especially characteristic of AI-generated errors. AI is trained to treat “code that does not raise an error” as what the user wants, so when it encounters a failure condition it tends to wrap the problem and continue rather than stop and report it. The code looks more robust, but in reality it has closed off its own channel for signaling that something is wrong.\nFilling values incorrectly can be just as dangerous as dropping them. One of the most painful silent errors I have encountered in practical data analysis arose here. Missing intervals were filled by interpolation during preprocessing, but the interpolation used values from future time points through two-sided interpolation. This leakage of future information is called look-ahead bias. It does not raise an error either. Instead, the data become contaminated from that point onward, and the entire analysis built on top of them loses its meaning. In practice, it causes predictive performance to be overestimated. In analyses that use interpolation, we begin by checking “which time point’s information was used to create each filled value.”\nThe way to catch the error is to count what was dropped. Replacing pass inside except with a counter and printing “Skipped N lines” at the end is enough to make the error stop being silent.\n5. Three Questions to Ask in 30 Seconds Before Trusting Code That Ran # This verification routine checks all three hiding places at once. You can begin even if you cannot yet read code, and it is worth learning before Python syntax.\nCompare a small input with a hand calculation. Give the code five values and compare the output with an answer you calculated mentally. An error that hides in 40,000 rows cannot hide in five. Count how many values entered the calculation. One call to len() is enough. If the number supplied differs from the number calculated, the discrepancy is the silent error confessing. Check whether the first and last values were included. Boundary errors occur at the ends. Checking only the beginning and end catches many of them. One of these is also the first check I actually perform when I receive a result.\nFigure 3. Ask these three questions before trusting code that runs.\n“How do I verify AI-generated code?” sounds like a large question, but the answer begins small: do not relax simply because there was no error. Spend 30 seconds examining these three places. This habit is the first muscle of the “ability to verify” discussed in the previous article.\nBut what should we do on a day when the code stops loudly? If you have been pasting the entire red error message into AI without reading it, the next article will explain how to spend 30 seconds reading that red text before asking AI.\nPython Fundamentals for the AI Era\nPrevious article: If AI Writes All the Code, Do We Still Need to Learn Python? This series grows out of the same concerns as the book Python Fundamentals for the AI Era, scheduled for publication in September.\nDisclosure of Interests and Responsibility\nThe author wrote the 2026 book Python Fundamentals for the AI Era, which addresses this topic.\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"August 21, 2026","externalUrl":null,"permalink":"/en/posts/2026-08-25-ai-code-quiet-errors/","section":"Writing","summary":"Code that finishes without an error is not verified code. This article explains three recurring hiding places for silent errors in AI-written Python—data-type assumptions, range boundaries, and swallowed exceptions—and a 30-second verification routine.","title":"The AI-Written Python Code Ran—and It Was Wrong — Three Places Silent Errors Hide","type":"posts"},{"content":"When AI can write code for us, what should people learn?\nI am developing my answer as a Korean-language book series called the AI ERA SERIES.\nAI ERA SERIES 01\nPython Fundamentals in the Age of AI Your first Python course, starting from zero experience\nWhen AI can write all the code, should we trust that code as it is? By practicing how to predict output before copying code, this book builds the fundamentals needed to read and verify AI-generated code.\nKorean-language edition scheduled for September 1, 2026 · Chapter example code will be published on GitHub with the book.\nAI ERA SERIES 02\nPython Data Visualization in the Age of AI From a single chart to dashboard deployment\nWhen AI can draw every chart, who decides what the chart should show? Beginning with the principle of choosing the one thing to communicate before drawing, this book covers the path from matplotlib to dashboard deployment.\nKorean-language edition scheduled for September 1, 2026 · Example code and dashboards will be published on GitHub with the book.\nPublication news and links to the example code will appear on this page on the release date. Errata for each book will also be maintained here.\n","date":"August 16, 2026","externalUrl":null,"permalink":"/en/books/","section":"Donghyun Lee’s Trustworthy AI Notes","summary":"AI ERA SERIES — Korean-language books about what and how to learn in the age of AI.","title":"Books","type":"page"},{"content":"In the previous article, we examined how MC Dropout and deep ensembles can estimate part of predictive uncertainty. Both approaches make repeated predictions for the same input and inspect how far apart the answers are.\nBut when AI encounters data it did not see during training, will those predictions automatically spread farther apart?\nIt would be convenient if they did, but there is no guarantee. An input can have all the expected columns and units and no missing values, yet contain a combination of season, region, and observation method that was almost absent from the training data. The model may still return an answer without an error message, while reporting uncertainty much like usual.\nFigure 1. An unfamiliar input and a model that knows to be cautious are not the same thing.\nThe method I used is not exempt from this question. In a 2024 study published in the Journal of Cleaner Production, I combined MC Dropout with ICNN to produce PM10 and PM2.5 predictions together with 95% uncertainty ranges.[1] The fact that the model produced uncertainty ranges, however, does not mean that it automatically detects every change in season, region, or observation method. My method also needs criteria for recreating unfamiliar conditions and validating performance again.\nThe 30-Second Takeaway # This article treats OOD as a perspective for warning that an individual input is unfamiliar, and dataset shift as a perspective for comparing a change between the distributions of training and operating environments. Changes in the inputs, outcome proportions, and input–outcome relationship are different problems. There is no guarantee that uncertainty will increase automatically. Before deployment, recreate expected changes, evaluate performance and uncertainty together, and define criteria for human review and revalidation. 1. OOD and Dataset Shift # A data distribution includes not just the average value, but also how often values occur, the ranges they occupy, and the combinations in which they appear.\nDataset shift, or distribution shift, is a state in which the pattern that generated the training data differs from the pattern generating operational data. Out-of-distribution (OOD) data usually concerns whether an operational input came from a distribution different from the chosen in-distribution baseline. Making that judgment requires a reference distribution, a detection score, and a threshold.[5] The boundary between these terms varies somewhat across studies; this article uses the following operational distinction.\nPerspective Primary question Unit Dataset shift Has the data-generating pattern changed between the training environment and the current environment? Comparison between two periods or datasets OOD detection Is this incoming input within the familiar range? An individual input or batch of inputs The two overlap but are not identical. A distribution shift can occur gradually as the proportion of certain conditions changes, even though each individual input appears familiar and no single case triggers a clear OOD alert.\nAn OOD alert is a signal of unfamiliarity, not a verdict that the prediction is wrong. The absence of an alert is not a certificate that the prediction is correct.\n2. What Changed? — Inputs, Proportions, or Relationships # Distribution shifts become easier to understand when classified by what has changed. In the table below, X is the input and Y is the outcome to be predicted.\nShift What changed? Question to ask Simplified example Covariate shift The distribution of X changes; the relationship between X and Y is assumed to remain the same Which input conditions now occur more frequently than before? A temperature–humidity combination that was rare during training appears frequently in operation Label-prior shift The proportions of classification outcomes Y change; the pattern of X within each outcome is assumed to remain the same Have the proportions of outcome categories, such as normal and caution, changed? The proportion of caution cases rises, while the feature pattern within caution cases remains the same Relationship shift\n(concept drift as used here) The conditional distribution of Y given the same X changes Does the same input now imply a different outcome? A change in the outcome definition or generating mechanism changes the meaning of the same observation Covariate shift has long been studied as the problem of the training sample and the real-world target population having different input distributions.[2] In classification, label shift assumes that the proportions of outcome categories change while the input pattern within each category remains stable.[3] The scope of concept drift varies across the literature. The relationship shift in this article refers to the case in which the conditional distribution of Y given the same X changes.[4]\nReal changes do not always fit neatly into one of these three boxes and may overlap. When someone says, “Drift was detected,” do not treat that sentence as identifying the cause. Work through the questions in the table to narrow down what changed.\nIn particular, input data alone cannot confirm the absence of the relationship shift described here. The shape of the inputs can remain the same while their relationship with the correct outcome changes. Such a shift often becomes visible only after actual outcomes arrive.\nFigure 2. Under the single label of distribution shift, inputs, outcome prevalence, and the relationship between inputs and outcomes can change in different ways.\n3. Uncertainty Does Not Increase Automatically # An AI model does not reflect, “This situation is outside my experience.” It produces a number according to learned computational rules. A conventional classification neural network is trained to separate the classes it learned, but the ability to recognize unfamiliar inputs does not appear automatically. Experiments by Hendrycks and Gimpel showed that neural networks can assign high softmax scores to misclassified and OOD examples, making it difficult to interpret that score alone as confidence.[5] Hein and colleagues used theory and image experiments to analyze how, under certain settings, ReLU-family classifiers can make highly confident predictions even in regions far from the training data.[6]\nWhat about MC Dropout and deep ensembles? These methods reveal disagreement among computational paths or independently trained models, but all may share the same training data, similar architectures, and the same prediction objective. If those shared elements create a blind spot, multiple predictions can converge in the same wrong direction. In experiments across image, text, ad-click, and genomics classification settings, Ovadia and colleagues compared several uncertainty-estimation methods under distribution shift. Both accuracy and uncertainty quality deteriorated as the shift increased. Deep ensembles were relatively strong on most metrics, but the researchers concluded that substantial room for improvement remained.[7]\nCalculating uncertainty through repeated predictions and having that uncertainty respond properly to unfamiliar data are separate validation problems.\nThe MC Dropout used in my 2024 particulate-matter study is not an exception to this principle. The paper shows results within that study’s setting.[1] It should not be extended into a claim of automatic warning across every region, period, and observation method.\nIn operation, we therefore need to separate three signals.\nSignal Question it answers What it cannot tell us on its own Input-shift or OOD score How atypical does the detector judge this input relative to the reference data? Whether the actual prediction is wrong Predictive uncertainty How much do the answers from models or computational paths vary? Whether the input is OOD, or whether the range matches real-world error Recent outcome-based report card Are actual performance, calibration, and interval quality being maintained? The outcome of a current input for which the ground truth has not yet arrived An OOD score is not a universal probability or common distance scale. It must be interpreted together with the detection method, reference data, and threshold used. Disagreement among the three signals matters most. If the input is unfamiliar but uncertainty is low, or the input appears familiar while recent errors are rising, that is a signal for human review and revalidation.\n4. Stress Testing Before Deployment # We cannot collect every unfamiliar input in the world in advance. We can, however, deliberately create foreseeable changes and test how the model responds. The predeployment checklist is as follows.\nSummarize the training range on one page.\nRecord the period, region or institution, equipment, units, preprocessing, missing-value handling, ranges of key variables, and outcome definition. This turns “similar to the training data” into a set of concrete comparison points.\nRecreate expected shifts through data splits.\nDo not rely only on random splits. Hold out an entire later period, an unseen region or institution, or different equipment and collection conditions for testing.\nTest several levels of severity for each shift type.\nThe severity scale is not a common unit for comparing different kinds of shift, however, and passing one OOD dataset does not prepare a model for every unfamiliar input.\nRead two report cards together.\nThe first covers predictive performance, such as accuracy or error. The second covers probability calibration and the coverage and width of prediction intervals. “Was it wrong?” and “Did it say ‘I don’t know’ properly?” are different questions.\nDefine the action that follows an alert.\nDecide at what level to request additional observations, require human review, or suspend automated prediction. There is no universal threshold because it depends on the costs of false alarms and missed detections and on the available review staff.\nPut revalidation triggers in the operational record.\nThese include a new season, region, or device; changes to preprocessing or the outcome definition; sustained input shifts; and deterioration in recent performance, calibration, or interval quality. A retrained model is a new model, so the same procedure must be repeated.\nFigure 3. OOD readiness is not one detection score; it combines shift scenarios, two scorecards, action rules, and revalidation triggers.\nFor a task in which ground truth arrives late, input-shift and OOD scores are early warnings; actual performance and uncertainty quality after outcomes accumulate form a delayed report card. Do not decide to retrain on an early warning alone. Examine what changed and confirm performance deterioration with recent outcomes. Conversely, even when the input alert is quiet, a failing outcome-based report card should trigger revalidation.\nConclusion # Does AI really say “I don’t know” when it encounters unfamiliar data? The answer is: you cannot know from the method’s name alone. MC Dropout, deep ensembles, and OOD detection are all useful, but none comes with a guarantee that it will work automatically when seasons, regions, or observation methods change in the field.\nI therefore trust an operational record that says, “We defined which changes count as unfamiliar, tested performance and uncertainty under those changes, and require human review and revalidation when the criteria are exceeded,” more than a description that says, “Our model knows when it does not know.”\nRelated Articles # How Does AI Calculate ‘I Don’t Know’? — MC Dropout vs. Deep Ensembles Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals Sources # [1] Lee, D., \u0026amp; Lee, B. (2024). Building reliable AI for quantifying uncertainty in particulate matter predictions with deep learning. Journal of Cleaner Production, 473, 143457. https://doi.org/10.1016/j.jclepro.2024.143457\n[2] Shimodaira, H. (2000). Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90(2), 227–244. https://doi.org/10.1016/S0378-3758(00)00115-4\n[3] Lipton, Z. C., Wang, Y.-X., \u0026amp; Smola, A. J. (2018). Detecting and Correcting for Label Shift with Black Box Predictors. Proceedings of the 35th International Conference on Machine Learning, PMLR 80, 3122–3130. https://proceedings.mlr.press/v80/lipton18a.html\n[4] Widmer, G., \u0026amp; Kubat, M. (1996). Learning in the Presence of Concept Drift and Hidden Contexts. Machine Learning, 23, 69–101. https://doi.org/10.1023/A:1018046501280\n[5] Hendrycks, D., \u0026amp; Gimpel, K. (2017). A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. International Conference on Learning Representations. https://openreview.net/forum?id=Hkg4TI9xl\n[6] Hein, M., Andriushchenko, M., \u0026amp; Bitterwolf, J. (2019). Why ReLU Networks Yield High-Confidence Predictions Far Away From the Training Data and How to Mitigate the Problem. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 41–50. https://openaccess.thecvf.com/content_CVPR_2019/html/Hein_Why_ReLU_Networks_Yield_High-Confidence_Predictions_Far_Away_From_the_CVPR_2019_paper.html\n[7] Ovadia, Y. et al. (2019). Can You Trust Your Model\u0026rsquo;s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. Advances in Neural Information Processing Systems 32. https://proceedings.neurips.cc/paper_files/paper/2019/hash/8558cb408c1d76621371888657d2eb1d-Abstract.html\nThe author provides related services in this field through AI Korea. This article does not promote any particular product or service.\nPredictions always contain uncertainty. They must not be used as the sole basis for decisions on disease-control or environmental policy.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"August 16, 2026","externalUrl":null,"permalink":"/en/posts/2026-08-16-dataset-shift-ood/","section":"Writing","summary":"AI uncertainty does not automatically increase when a model encounters data unlike its training data. This article distinguishes OOD from three kinds of distribution shift and sets out criteria for stress testing and revalidation.","title":"Does AI Really Say ‘I Don’t Know’ When It Encounters Unfamiliar Data? — Dataset Shift and OOD","type":"posts"},{"content":"Imagine the moment when a new infectious disease has just begun to spread. We need an AI forecasting model to anticipate how the outbreak will develop. Yet we have only a few weeks of domestic data with which to train it.\nOne method for this situation is transfer learning. We first train a model on data already accumulated in another country, then adapt it to our own data.\nThat raises the next question: Which country should the model learn from?\nChoosing a country similar to ours seems obvious. Experience from a country with a similar epidemic curve would appear more likely to transfer well.\nWhen I tested that intuition across 30 countries, the opposite was true.\nFigure 1. In early outbreak forecasting without local data, the first problem is deciding whom to learn from.\nThis article explains for a general audience a paper that I published as sole author in Expert Systems with Applications (IF 9.4).[1] It is also a preview of the “What It Means to Forecast Infectious Disease” series that will begin this fall.\nThe 30-Second Takeaway # When training an infectious-disease forecasting model on data from other countries, models that learned from the most “different” countries consistently outperformed those that learned from the most “similar” countries. The benefit did not increase in proportion to dissimilarity. A statistically significant improvement appeared only in the most heterogeneous quartile. This is an option for the early stage of an emerging infectious disease, when domestic data are unavailable. Its limited validation conditions, however, must be read alongside the result. 1. Why Do We Need Data from Other Countries? — The Data Gap at the Beginning of an Emerging Outbreak # Deep-learning models grow by consuming data. With enough data they can be powerful, but without data, even an excellent architecture has nothing to learn.\nThe beginning of an emerging infectious disease is exactly that situation. The moment when forecasts are most urgently needed coincides with the moment when data are scarcest. When building infectious-disease forecasting models, I find this shortage of early data especially frustrating. When a new disease emerges, domestic data about that disease have not yet accumulated.\nTransfer learning is one way to fill the gap. It resembles the way a person who has learned one foreign language can learn a second more quickly. Instead of beginning from the idea of a sentence itself, the learner starts with an existing sense of “how languages generally work.” A model likewise learns first from outbreak data in other countries, then completes its training on domestic data.\nThe problem is how to choose that “other country.” Most intuitions point in one direction: a country similar to ours. I was often asked the same question: “Surely transfer learning should use data similar to ours?”\nThis study tested that “surely” against the data.\n2. A Similar Country Seems Like the Obvious Choice — The Opposite Result from 30 Countries # The experiment was designed as follows. I used 1,143 days of publicly available COVID-19 data from 30 countries, covering January 22, 2020, through March 9, 2023. The sources were the Johns Hopkins University aggregation (JHU CSSE) and Our World in Data (OWID).[2][3]\nFor each target country, I selected four source countries, pretrained a model on their data, and then fine-tuned it on the target country’s data. Similarity between countries was measured with a distance called dynamic time warping (DTW), which measures how closely the shapes of two time-series curves resemble one another.\nI compared several source-selection strategies: the four countries most similar to the target, the four most dissimilar, four selected at random, and an individual model trained only on the target country’s own data without transfer learning. The model itself was held constant: one lightweight time-series neural network called a TCN.\nHere is the result.\nThe model that learned from the most dissimilar countries—heterogeneous transfer learning—reduced prediction error (RMSE) by about 40% compared with the individual model trained only on domestic data (52.5 → 31.5). RMSE measures how far predictions deviate from observed values on average; lower is better.\nThe relative metrics told the same story. RMSE normalized by the standard deviation of the observed data (RMSE/SD) was 0.37, lower than the 0.62–0.75 of the individual models. R², which indicates how much of the observed variation was explained, was 0.82.\nI also tested whether these differences might be due to chance. Using the Wilcoxon signed-rank test to pair the errors of the two methods in each of the 30 countries, the differences from all individual models had p\u0026lt;.001; the difference from transfer learning with randomly selected sources had p=.009; and the difference from transfer learning with similar-country sources also had p\u0026lt;.001. A p-value is a scale for how likely a difference of this size would be to arise by chance. By convention, a value below 0.05 is treated as difficult to attribute to chance.\nIn short, learning from similar countries was not a bad choice, but learning from the most dissimilar countries was consistently better. The direction ran against conventional wisdom.\n3. A Stranger Finding — The Benefit Appeared Only at the “Most Dissimilar” Extreme # At this point, it might seem that we can simply replace the old intuition with a new one: “the more different, the better.” The data did not permit that conclusion either.\nI divided source countries into four quartiles based on their dissimilarity from the target country. Q1 was the most similar quartile, and Q4 the most dissimilar. Average prediction error (RMSE) across the 30 countries was as follows.\nQuartile Q1 (most similar) Q2 Q3 Q4 (most dissimilar) Mean RMSE 35.29 34.26 34.21 31.00 Q1, Q2, and Q3 were not statistically distinguishable from one another. Only Q4 had an error 12.1% lower than Q1, and only this difference was significant (p=.008).\nFigure 2. After source countries were grouped by heterogeneity, a significant error reduction appeared only in the most heterogeneous quartile, Q4. The vertical axis begins at zero.\nThe chart’s vertical axis begins at zero. The decline therefore does not look dramatic, and that is the correct impression. Truncating the axis to inflate a difference is precisely the habit this blog opposes. The difference is modest but statistically clear.\nThe strangest and most interesting aspect of this study is that the benefit did not rise gradually along a similarity scale; it was concentrated at the extreme. Being somewhat different did not help. The benefit appeared only when the data were very different.\nThe more a result contradicts intuition, the more important it becomes to check it against data rather than intuition.\n4. My Interpretation — Similar Countries Learn Each Other’s “Dialect” # Here I need to separate fact from interpretation. The evidence established by the paper ends with Section 3. The paper does not yet explain the mechanism that concentrates the benefit in the most heterogeneous quartile.\nMy interpretation is as follows.\nImagine someone trying to learn a standard language from speakers in only one region. They may learn the region’s dialect as though it were part of the standard language. A model trained only on similar countries may do something comparable. It may memorize local habits those countries happen to share—aggregation methods, reporting cycles, or the stage at which an outbreak passed—as though they were general laws of transmission. When sufficiently heterogeneous data are combined, by contrast, only the common grammar of transmission that survives all those differences remains.\nThis account fits the result well, but it is a hypothesis consistent with the result, not a mechanism demonstrated by the paper. That is where the next study begins.\nTwo practical points are worth adding. First, this method is an option that can actually be used early in an emerging outbreak, before domestic data have accumulated. Second, the result was achieved not with a heavy foundation model, but with one lightweight TCN. In settings with limited compute, that difference is not trivial.\n5. Reading the Limitations as Written — The Same Standard Applies to the Models I Build # It is time to say how far this result can be trusted. The limitations are clear.\nFirst, the evaluation covered one low-incidence period late in the pandemic. The study did not test whether the same conclusion would hold during a sharp surge in cases.\nSecond, it evaluated one-day-ahead forecasts only. The paper does not answer whether the benefit of heterogeneous transfer persists for forecasts one week or one month ahead.\nThird, as noted above, it did not explain why the benefit was concentrated at the extreme. When we do not know why a method works, it is difficult to predict when it will stop working as conditions change.\nFigure 3. The study evaluated one late-pandemic, low-incidence period at a one-day forecast horizon; the result must be read with those conditions attached.\nThese three questions apply directly to my own work as well. I build infectious-disease forecasting models not only in papers, but also for field settings. What period of data was used for validation? How many steps ahead were tested? Can we explain why the model works? The fact that this paper has been held to these standards does not mean that other forecasting models I build in practice pass them automatically. They must face the same questions in the same way.\nWhat I hope you take from this article, then, is not one number but one standard for judgment:\nValidate the data choices that seem “obvious.” Intuition is a hypothesis, not evidence.\nFigure 4. Intuition is a hypothesis, not evidence; the more obvious a choice appears, the more it needs testing.\nConclusion # A model learned more effectively from the most dissimilar countries than from the most similar ones; the benefit appeared only at the extreme; and the result was established under the conditions of a low-incidence period and one-day-ahead forecasting. I hope you will remember the result and its conditions in the same sentence.\nBeginning this fall, the “What It Means to Forecast Infectious Disease” series will examine this question in earnest. Timed to the avian-influenza season, it will explore what evidence can support forecasts when data are scarce, and when those forecasts should be trusted or questioned.\nDoes your field have a choice that seems so obvious that no one has ever tested it?\nRelated Articles # How Does AI Calculate ‘I Don’t Know’? — MC Dropout vs. Deep Ensembles Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals Sources # [1] Lee, D. (2027). Heterogeneous transfer learning for robust infectious disease forecasting: A data-centric approach. Expert Systems with Applications, 332, 133728. https://doi.org/10.1016/j.eswa.2026.133728\n[2] Johns Hopkins University CSSE. COVID-19 Data Repository. https://github.com/CSSEGISandData/COVID-19\n[3] Our World in Data. Coronavirus Pandemic (COVID-19) Data. https://ourworldindata.org/coronavirus\nAll figures reported in the body of this article (RMSE 52.5→31.5, RMSE/SD 0.37, R² 0.82, quartile values 35.29/34.26/34.21/31.00, 12.1%, and the p-values) come from the published version of [1].\nThe author provides related services in this field through AI Korea. This article does not promote any particular product or service.\nPredictions always contain uncertainty. They must not be used as the sole basis for decisions on disease-control or environmental policy.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"August 11, 2026","externalUrl":null,"permalink":"/en/posts/2026-08-11-heterogeneous-transfer-learning/","section":"Writing","summary":"In transfer learning for infectious-disease forecasting, models trained on data from the most dissimilar countries consistently outperformed those trained on the most similar countries. This article explains the 30-country experiment and the limitations that must accompany its results.","title":"When a New Infectious Disease Has Little Data, Which Country Should AI Learn From? — Infectious Disease Forecasting and Heterogeneous Transfer Learning","type":"posts"},{"content":"Donghyun Lee (the “Operator”) considers the personal information of donghyunlee.kr users important and publishes this Privacy Policy in accordance with Korea’s Personal Information Protection Act and other applicable laws.\nThis policy applies to donghyunlee.kr. The site does not provide user registration, comments, payments, or an inquiry form.\n1. Personal information processed and purposes # Site usage analytics # The Operator uses Google Analytics 4 to understand site usage and improve the content and user experience. Related information is collected automatically when a user accesses the site. Users who do not want this collection may opt out using the methods in Section 5 below.\nData processed: a randomly generated client identifier; pages and page titles viewed; dates and times of visits and interactions; referral source; browser, operating system, device, and screen information; and approximate location Purpose: compiling visitor statistics, analyzing content usage, and improving the site Method: automatic collection through the Google tag and first-party cookies (_ga, _ga_\u0026lt;measurement ID\u0026gt;) when the site is accessed Google states that Google Analytics may use an IP address during network communication and to determine approximate location, but does not log or store individual IP addresses. The Operator does not send Google Analytics information that directly identifies a person, such as a name, email address, or telephone number.\nEmail inquiries # When a user directly contacts the email address displayed on the site, the sender’s email address, name, affiliation, and inquiry may be processed. This information is used only to review and answer the inquiry. Please do not include sensitive or unnecessary personal information in an email.\nIf an inquiry concerns laboratory or company work, the Operator will direct the user to the appropriate official contact channel and will not forward the inquiry to another entity without the user’s separate consent.\n2. Processing and retention periods # Google Analytics user- and event-level data: 14 months from collection Google Analytics cookies: up to 14 months from creation; the expiration period is not extended with each visit Email inquiry information: deleted without delay after the purpose of the inquiry has been fulfilled, unless retention is required for a contract, dispute, or under applicable law Statistics that are not directly tied to a personal identifier, such as standard aggregated Google Analytics reports, may be retained longer under Google’s policies.\n3. Google Analytics and processing outside Korea # When a user accesses the site, information may be processed outside Korea as follows:\nRecipient: Google LLC (1600 Amphitheatre Parkway, Mountain View, CA 94043, USA) Country: United States Data transferred: the data listed under “Site usage analytics” above Timing and method: transmitted through an encrypted network while the site is being used Purpose: providing Google Analytics, processing visitor statistics, and generating reports Retention and use period: 14 months for user- and event-level data Opt-out method and effect: users may opt out using the methods in Section 5. Opting out does not limit use of the site. For information on how Google processes data, see How Google uses information from sites or apps that use its services.\n4. Provision to third parties and outsourced processing # The Operator does not provide users’ personal information to third parties beyond its original purpose, except with the user’s consent or where specifically permitted by law.\nSite analytics are processed through Google Analytics, provided by Google LLC. Google processes analytics data according to the Operator’s settings and the contracts and policies that apply to Google Analytics.\n5. How to opt out of collection # Users may opt out of Google Analytics data collection in any of the following ways. None of these methods limits use of the site.\nBlock or delete this site’s cookies in browser settings Use the browser’s tracking-prevention feature or private/incognito mode Install the Google Analytics Opt-out Browser Add-on 6. User rights and how to exercise them # Users may request access to, correction or deletion of, suspension of processing of, or withdrawal of consent regarding their personal information. Please send requests to the email address below. When identity verification is necessary, the Operator may request the minimum information required and will provide the result in accordance with applicable law.\n7. Personal information protection and inquiries # Personal information controller and contact: Donghyun Lee Email: donghyunlee.ai@gmail.com The Operator limits access to personal information to what is necessary and maintains security measures available for the site and email account.\n8. Changes to this policy # If this policy changes, the changes and their effective date will be published on this page.\nEffective date: July 27, 2026 Last revised: August 24, 2026 (Google Analytics retention changed from two months to 14 months; cookie lifetime changed from 60 days to a maximum of 14 months; the consent-banner approach changed to notice and opt-out guidance.) ","date":"July 27, 2026","externalUrl":null,"permalink":"/en/privacy/","section":"Donghyun Lee’s Trustworthy AI Notes","summary":"How donghyunlee.kr processes personal information and uses Google Analytics.","title":"Privacy Policy","type":"page"},{"content":"I fed the same particulate-matter data into the same AI model. Yet the result changed slightly each time I made a prediction.\nIs the model broken?\nWith MC Dropout, we create this variation deliberately. If one run returns 41, the next 44, and another 39, we examine how widely those values are spread to estimate the model’s uncertainty.\nI used this method in actual research on particulate-matter prediction. In a 2024 paper published in the Journal of Cleaner Production, I combined MC Dropout with a deep-learning model called ICNN to produce PM10 and PM2.5 predictions together with uncertainty ranges.[5]\nThat does not mean MC Dropout is always the best method. Deep ensembles are also widely used for a similar purpose. Both approaches collect multiple predictions, but the way they generate those predictions is completely different.\nFigure 1. Both methods produce multiple predictions, but they vary different parts of the process.\nThe 30-Second Takeaway # Bayesian deep learning is a broad framework; MC Dropout is one way to approximate it. MC Dropout changes paths within one model, while a deep ensemble trains multiple models separately to create variation among predictions. The fact that predictions vary does not mean that their variation honestly reflects real-world error. [1] First, Correct the Map of the Comparison # We can deliberately produce different answers for the same input.\nWe ordinarily expect the same model to return the same answer for the same input. To estimate uncertainty, we deliberately disturb this rule. We examine multiple possible models or paths and observe how tightly their answers cluster or how widely they spread.\nA conventional deep-learning model learns one set of the most plausible weights.\nA neural network adjusts a large number of weights during training. In ordinary prediction, it uses one set of trained weights to produce one answer. The answers that other possible models might have produced do not appear on the screen.\nThere is more than one reason for AI not to know.\nUncertainty inherent in the data, such as measurement error or random variation, is called aleatoric uncertainty. When the model does not know because training data are limited or conditions are unfamiliar, the uncertainty is closer to epistemic uncertainty.[1] The two cannot always be separated perfectly in practice, but differences among MC Dropout or deep-ensemble predictions are used mainly as clues to uncertainty on the model side.\nBayesian deep learning and MC Dropout are not competing alternatives.\nBayesian deep learning is a broad approach that represents multiple plausible weights or functions as a distribution in light of the data, rather than fixing one set of weights or one function. MC Dropout was proposed as a computationally tractable way to approximate this Bayesian inference using dropout neural networks.[2] Treating the two as an A versus B choice therefore compares a broad category with one method inside it.\nA deep ensemble is a more natural comparison with MC Dropout.\nA deep ensemble also creates multiple predictions and examines their differences to study uncertainty. Rather than approximating an explicit posterior distribution, however, it trains several neural networks independently from the beginning. The original paper presented this as a practical alternative to approximate Bayesian neural networks.[3]\n[2] MC Dropout Creates Multiple Opinions Within One Model # Dropout was originally a training method for reducing overfitting.\nDuring training, it randomly makes some units in a neural network sit out. This prevents the model from always depending on the same combination of units. It resembles having different groups of people take turns solving a problem so that all the work does not fall on one person.[4]\nOrdinarily, dropout is turned off at prediction time.\nSome units are randomly dropped during training, but conventional inference turns dropout off and uses one stable computational path. That is why the same input produces the same answer.\nMC Dropout keeps dropout on during prediction.\nMC stands for Monte Carlo, a method that approximates a quantity by drawing random samples repeatedly. Because a different combination of units sits out on each prediction, the process has the effect of sampling slightly different virtual models from within one model.\nThe average of the predictions becomes the answer; their spread offers a clue to what the model does not know.\nPassing the same input through the model repeatedly produces a collection of predicted values. Their mean or median can serve as the representative prediction, while their dispersion can serve as an uncertainty measure. More repetitions reduce the variability of this sampling calculation, but they do not eliminate bias or flawed assumptions in the original model.\nI used this structure to predict particulate matter.\nIn the 2024 paper, we arranged air-quality and meteorological data in a multidimensional grid and combined ICNN with MC Dropout to produce PM10 and PM2.5 concentrations together with 95% uncertainty ranges.[5] The goal was to show not only a single prediction, but also how much it might vary. The use of MC Dropout in this study does not mean it is best for every region, time period, or environmental model. The verification criteria described in this article apply equally to the method I used.\n[3] Deep Ensembles Gather the Opinions of Multiple Models # A deep ensemble trains multiple models separately.\nA typical deep ensemble begins several neural networks with the same architecture from different initial values and optimizes each one separately. Even with the same training data, differences in starting points, minibatch order, and other factors can lead to slightly different final weights.\nThe multiple models resemble different teams solving the same problem.\nIf MC Dropout is like repeatedly asking one team while changing some of its members, a deep ensemble is more like training several teams independently and then asking each the same question. Similar answers converge on one direction; substantially different answers signal greater uncertainty arising from model selection.\nLook not only at the average, but also at the disagreement among models.\nAveraging the predictions from several models creates one final prediction. At the same time, we can calculate how far apart their answers are. Because all the models may share similar data, architectures, and blind spots, however, agreement does not guarantee that the answer is correct.\nFigure 2. MC Dropout changes paths inside one trained model; a deep ensemble repeats the training itself.\nDeep ensembles cost more to train and store.\nIf an ensemble contains five models, it typically requires five training runs and five stored models. MC Dropout has the advantage of training and storing one model. At prediction time, however, MC Dropout also requires multiple forward passes, so “one model” does not mean “one inference.” Both approaches must be assessed in light of parallel processing and response-time requirements.\nDeep ensembles are a strong baseline, but not a universal answer.\nIn the classification and regression experiments in the original deep-ensemble paper, their uncertainty quality was similar to or better than that of approximate Bayesian neural networks, and the authors also presented experiments showing greater uncertainty on out-of-distribution inputs.[3] This does not prove that deep ensembles are always better than MC Dropout. Rankings can change with the data, architecture, training budget, and evaluation method.\n[4] After Choosing a Method, Verify the Uncertainty Again # The difference between the two methods can be reduced to this table.\nComparison MC Dropout Deep Ensemble Typical training One model Multiple models Storage One set of weights Multiple sets of weights Repeated prediction Multiple runs with changing dropout masks One or more runs per model Source of diversity Different paths within one model Multiple independently trained models Representative advantage Relatively easy to apply to an existing dropout architecture Strong empirical baseline Representative burden Repeated inference; dependence on dropout design Training and storage cost There is no universally correct number of repetitions or models. It should be selected by considering task accuracy, latency, and uncertainty quality together.\nPredictions that vary are not necessarily predictions whose variation is honest.\nA model does not automatically capture real-world error well simply because it produces a wide range. Conversely, a narrow range does not prove genuine certainty. Producing uncertainty and checking whether that uncertainty matches observed frequencies—calibration—are separate tasks.\nCheck whether the model really becomes more cautious on unfamiliar data.\nChanges in season, region, or observation equipment can introduce data unlike those seen in training. In a large comparative study, as such distribution shifts grew, both accuracy and uncertainty estimation and calibration could deteriorate.[6] A recent study also reported cases in its experimental setting where MC Dropout did not adequately reflect increased uncertainty in interpolation and extrapolation regions.[7] One study cannot invalidate all uses of MC Dropout, but neither should we assume that the model will “automatically become anxious when the input is unfamiliar.”\nA model needs two report cards.\nThe first reports how accurate its predicted values are. The second reports how often the uncertainty interval contains the actual value and whether the interval is excessively wide. A model can have a small prediction error while understating uncertainty, or it can increase coverage by producing ranges too wide to be useful.\nWrite down the selection criteria before the method name.\nIf the existing model uses dropout and you need to establish a baseline quickly, MC Dropout may be practical. If you have the resources to train and store multiple models and need a strong comparison, you can test a deep ensemble as well. For important decisions, do not choose either method by name alone. Under the same data and compute budget, compare prediction error, interval coverage and width, distribution shift, and response time.\nFigure 3. Choosing a method is only the start; its uncertainty quality must be tested on realistic data alongside its computational cost.\nConclusion # MC Dropout and deep ensembles both look for what AI does not know in the differences among multiple predictions.\nMC Dropout runs one model in multiple forms. A deep ensemble trains multiple models separately. The former has lower training and storage costs; the latter gathers independently trained outcomes at greater cost.\nThe most important distinction, however, does not lie in the method names.\nAI’s “I don’t know” can be calculated from disagreement among predictions. But that disagreement earns the right to be trusted only through validation on real-world data.\nRelated Articles # Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals Why Do We Trust Claims More Easily When They Include Numbers? — 20 Short Thoughts on Reading Numbers Sources # [1] Kendall, A., \u0026amp; Gal, Y. (2017). What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision? Advances in Neural Information Processing Systems 30. https://papers.nips.cc/paper_files/paper/2017/hash/2650d6089a6d640c5e85b2b88265dc2b-Abstract.html\n[2] Gal, Y., \u0026amp; Ghahramani, Z. (2016). Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. Proceedings of the 33rd International Conference on Machine Learning, PMLR 48, 1050–1059. https://proceedings.mlr.press/v48/gal16.html\n[3] Lakshminarayanan, B., Pritzel, A., \u0026amp; Blundell, C. (2017). Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. Advances in Neural Information Processing Systems 30. https://papers.nips.cc/paper_files/paper/2017/hash/9ef2ed4b7fd2c810847ffa5fa85bce38-Abstract.html\n[4] Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., \u0026amp; Salakhutdinov, R. (2014). Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research, 15, 1929–1958. https://jmlr.org/papers/v15/srivastava14a.html\n[5] Lee, D., \u0026amp; Lee, B. (2024). Building reliable AI for quantifying uncertainty in particulate matter predictions with deep learning. Journal of Cleaner Production, 473, 143457. https://doi.org/10.1016/j.jclepro.2024.143457\n[6] Ovadia, Y. et al. (2019). Can You Trust Your Model\u0026rsquo;s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. Advances in Neural Information Processing Systems 32. https://proceedings.neurips.cc/paper_files/paper/2019/hash/8558cb408c1d76621371888657d2eb1d-Abstract.html\n[7] Djupskås, A., Riemer-Sørensen, S., \u0026amp; Stasik, A. J. (2026). Unreliable Monte Carlo Dropout Uncertainty Estimation. Proceedings of the 7th Northern Lights Deep Learning Conference, PMLR 307, 106–114. https://proceedings.mlr.press/v307/djupskas26a.html\nThe author provides related services in this field through AI Korea. This article does not promote any particular product or service.\nPredictions always contain uncertainty. They must not be used as the sole basis for decisions on disease-control or environmental policy.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"July 26, 2026","externalUrl":null,"permalink":"/en/posts/2026-07-26-mc-dropout-vs-deep-ensembles/","section":"Writing","summary":"MC Dropout runs one model in many forms; deep ensembles train multiple models separately. Twenty accessible points explain how these two approaches calculate what AI does not know, along with their costs and limitations.","title":"How Does AI Calculate ‘I Don’t Know’? — MC Dropout vs. Deep Ensembles","type":"posts"},{"content":"The following two statements are hypothetical examples created for illustration.\nUser satisfaction increased after the new approach was introduced.\nUser satisfaction was 73.6% after the new approach was introduced.\nThe second statement may make us look twice. It sounds as though a questionnaire exists somewhere, along with a spreadsheet and someone who checked the result. The first statement sounds like an opinion; the second sounds like evidence.\nYet we have not verified anything. We do not know how many people were asked, who they were, what the question was, or when the survey took place. The only new information is the presence of a decimal point.\nNumbers are a powerful language for comparing reality and making decisions. Problems do not arise only when a number is wrong. They also arise when a number looks so convincing that we stop asking questions.\nFigure 1. A decimal point can make a sentence look like a measurement rather than an opinion.\nThe 30-Second Takeaway # Numbers reduce ambiguity, but the number of decimal places does not guarantee the quality of the evidence. Every number contains human choices about definitions, denominators, time periods, averages, and comparison points. To trust numbers well, look beyond the value to the path that produced it, its uncertainty, and the action it will change. [1] Numbers Earn Trust Before They Earn the Status of Fact # “It improved by 37.2%” can feel more trustworthy than “It improved a lot.”\nThe 37.2% is hypothetical. Even so, it can sound more like a solid measurement than “a lot” because an ambiguous adjective has been converted into a comparable value.\nBut we still have not seen the raw data or the formula.\nThe existence of a number is not the same as the existence of good evidence behind it. Being calculated is not the same as being calculated properly.\nA precise number can be read as a signal that “someone knows the details.”\nExperimental research has found conditions in which precise numbers influenced subsequent estimates more strongly than round numbers. The effect did not always appear, however. It emerged when people assumed that the speaker had relevant knowledge and a reason for communicating that precision, and when the level of detail seemed necessary for the task.[1]\nDecimal places alone cannot tell us whether a value is accurate.\nIn metrology, “precision” refers to how close repeated measurements of the same object are to one another, while “accuracy” refers to how close a measured value is to the true value. Repeated measurements can be close to one another without being close to the truth.[2] If a thermometer consistently reads 0.5 degrees too high under the same conditions, its measurements may be consistent but not accurate.\nAn opinion reveals the person speaking; a number can appear to speak for itself.\nThe phrase “In my view” makes the interpreter visible. In a table or graph, however, the people who chose the question and selected the data can easily disappear into the background. This is one reason numbers look objective.\n[2] Numbers Do Not Remove Judgment; They Move It Upstream # Every number begins with a definition of what to count.\nDoes a “satisfied respondent” mean someone who chose 4 or 5 on a five-point scale, or only someone who chose 5? Change the definition, and the satisfaction rate changes even with the same responses.\nEvery proportion has a denominator: “out of what?”\n“Eight out of ten” and “80% of respondents” may mean the same thing—or something entirely different. Whether nonrespondents can be excluded is a question for the survey design, not the number itself.\nA large sample is not sufficient by itself.\nEven a survey of 10,000 people can miss the answer if it includes only a group unlike the people we want to understand. Before looking at the size of the number, ask whom it represents.\nAverages are excellent at erasing differences.\nTwo groups can have the same average even when one is clustered near the middle and the other is split between the extremes. An average is a useful summary, but it does not preserve the original shape of the data.\nThe analytical method is part of the number.\nIn one study, 29 analysis teams examined the same data and the same question, yet their effect estimates and conclusions varied substantially depending on choices about variables and models.[3] This does not mean statistics can say anything we want. It means that we must disclose the choices that produced a number before its meaning can be judged.\nFigure 2. Judgments omitted from a number do not disappear; they move upstream into how the number is made.\n[3] The Same Number Changes Meaning with Its Context # Every rate of change has a starting point.\nAn increase from one visitor to two is a 100% increase. “It doubled” is true, and “it increased by one person” is also true. What we choose to show alongside the number changes how large the increase feels.\nThe comparison point sits outside the number but supplies half its meaning.\nA score of 20 tells us nothing about whether the result is good or bad. Judgment begins only when we know whether yesterday’s score was 10, the target is 100, or a comparable group averages 18.\nChanging the period can make the same phenomenon look like a different story.\nA one-day surge may look like a crisis, while the same movement may look small in a one-year trend. Neither the short nor the long period is always correct. What matters is whether the period fits the decision at hand.\nMathematically equivalent outcomes can produce different choices depending on how they are framed.\nClassic research on decision-making found that people made different choices when the same outcomes were presented as gains rather than losses.[4] Numbers do not operate outside language and context.\nThe first number we see can easily become the starting point for the next judgment.\nThe “anchoring effect,” in which a previously presented number influences a later estimate under uncertainty, has been studied for decades.[5] But not every number has the same influence in every situation. One practical response is to write down your own criterion before seeing the first number.\n[4] A Good Number Reports Its Own Limits # Past average performance is not the certainty of this one case today.\nThis is a gap I repeatedly encounter while researching infectious-disease and environmental prediction models and considering their use in the field. It is not enough to say that a model performed well overall on past data. What people in the field want to know is how much they can trust the prediction in front of them now, and what they can use it for.\nThe value of a number appears in the action it changes.\nA number that is merely reported and forgotten should not require the same level of evidence as one that moves people, budgets, or time. The more important the decision, the more we need to consider not only the value but also the cost of being wrong and the conditions for reversing the decision.\nDisclosing uncertainty does not necessarily destroy trust.\nFour online experiments and a field experiment on the BBC News website included 5,780 participants in total. Across the studies, communicating uncertainty reduced trust in the number itself, but the decline was small when uncertainty was presented as a numerical range. Trust in the source barely declined with numerical ranges; larger declines occurred with vague verbal expressions.[6] One study cannot be generalized to every setting, but honestly reporting a range does not necessarily make trust collapse.\nRecovering just five things can change the story a number tells.\nWhat exactly was counted? — Definition Out of what was it calculated? — Denominator What was it compared with? — Comparison How much could it vary? — Uncertainty What decision will it inform? — Use A number is the beginning of a question, not the end of one.\nThis is not an argument against trusting statements that contain numbers. It is an argument against ending our judgment merely because a number is present. A good number makes what we know clearer while revealing what we do not know as well.\nFigure 3. The best way to question a number is not to reject it, but to reconstruct the path by which it was made.\nConclusion # Numbers make complex realities visible at a glance. That is why we need them. At the same time, the process of making reality visible at a glance folds away a great deal of context. That is why we need questions.\nThe next time you encounter a solid-looking number such as 73.6%, do not begin with the decimal places. Check five things first:\nDefinition, denominator, comparison, uncertainty, and use.\nThe moment a number looks precise is not the end of our judgment. It is where judgment should begin.\nRelated Articles # Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals If AI Writes All the Code, Do We Still Need to Learn Python? — From Writing to Verification Sources # [1] Zhang, Y. C., \u0026amp; Schwarz, N. (2013). The power of precise numbers: A conversational logic analysis. Journal of Experimental Social Psychology, 49(5), 944–946. https://doi.org/10.1016/j.jesp.2013.04.002\n[2] Joint Committee for Guides in Metrology. International Vocabulary of Metrology — Measurement accuracy (2.13); Measurement precision (2.15). https://jcgm.bipm.org/vim/en/2.13.html · https://jcgm.bipm.org/vim/en/2.15.html\n[3] Silberzahn, R. et al. (2018). Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results. Advances in Methods and Practices in Psychological Science, 1(3), 337–356. https://doi.org/10.1177/2515245917747646\n[4] Tversky, A., \u0026amp; Kahneman, D. (1981). The Framing of Decisions and the Psychology of Choice. Science, 211(4481), 453–458. https://doi.org/10.1126/science.7455683\n[5] Tversky, A., \u0026amp; Kahneman, D. (1974). Judgment under Uncertainty: Heuristics and Biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124\n[6] van der Bles, A. M. et al. (2020). The effects of communicating uncertainty on public trust in facts and numbers. Proceedings of the National Academy of Sciences, 117(14), 7672–7683. https://doi.org/10.1073/pnas.1913678117\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"July 25, 2026","externalUrl":null,"permalink":"/en/posts/2026-07-25-why-we-trust-numbers/","section":"Writing","summary":"Why does a claim with a decimal point feel more objective? Twenty short thoughts on recovering the definitions, denominators, comparisons, uncertainty, and uses hidden behind a number.","title":"Why Do We Trust Claims More Easily When They Include Numbers? — 20 Short Thoughts on Reading Numbers","type":"posts"},{"content":"How much will our work change by 2027? It is a question that appears in almost every talk about AI, but I think focusing on that question alone may cause us to miss a more important change.\nI often encounter the same scene in evaluations today. The analysis code runs without errors, but when I ask, “Why did you handle it this way?” the person cannot answer. AI wrote the code, and the person submitting it does not know what it does—a decidedly dangerous situation.\nBy 2027, the same scene may spread to reports, data analysis, research design, and business plans: an abundance of output, with no one who understands it or can take responsibility for it. The first risk that comes to my mind is not the disappearance of jobs, but this scene. I believe the opportunity lies in the same place. It is likely to go not merely to people who can extract answers from AI, but to those who can design work structures and evaluation criteria that make AI reason more deeply.\nOpening ChatGPT, Gemini, or Claude on the web or in an app, entering a prompt, and receiving an answer—this has been our basic way of using AI. If you think this question-and-answer pattern is enough preparation for the AI era, it is worth pausing for a moment.\nWhat is now emerging is a structure in which AI develops its own reasoning chain and work process. I am convinced that the gap between people who understand this structure and those who do not will grow not linearly but quadratically. This article is about understanding and learning that structure.\nFigure 1. The 2027 divide may depend less on whether people use AI than on who designs the structure in which AI works.\nThe 30-Second Takeaway # In 2027, AI will likely resemble a work environment that performs multistep tasks more than a tool that simply answers questions. Bundles of tasks and entry-level pathways will change before occupations do, while scarcity will shift from generation to goal setting, verification, and responsibility. The preparation that matters is not learning to use one particular model, but building five assets: domain knowledge, questions, harnesses, verification, and trust. 1. The Future Should Be Read as Scenarios with Different Probabilities, Not One Number # Let me draw a clear line first. No one can predict AI in 2027 precisely. I therefore divide the future into three layers. Relatively certain: AI will become more deeply embedded in documents, code, search, and analytical tools, and most knowledge workers will work with it. Likely, but uncertain in speed: AI will expand into agents that use multiple tools and carry out longer processes. Difficult to assert: the date on which a particular occupation will disappear or artificial general intelligence will be completed.\nAccording to Stanford’s 2026 AI Index, AI adoption among surveyed organizations reached 88%, but actual deployment of AI agents remained in the single digits across most business functions.[1] AI use has spread rapidly, but the stage at which it reshapes an organization’s work as a whole is only beginning. The forecasts in this article are therefore not declarations. They are a record of what is likely to change first if today’s signals continue—and what signals would tell us that the forecast was wrong.\n2. AI Will Move from a Prompting Tool to a “Thinking Harness” # Many people currently use AI as a question box: one question, one answer. By 2027, that approach alone is likely to be insufficient. What will matter more is the structure around the model—a framework connecting goals, context and memory, tools, human verification points, evaluation rubrics, records, and improvement. I will call this a thinking harness.\nAsking “Write a report” once is prompt use. Breaking the question down, gathering trustworthy sources first, searching for counterevidence to each claim, and independently checking calculations again is a harness. Add a rubric that lists the requirements for a good result, and it becomes a learning structure: weaknesses found in one result can be added to the evaluation criteria so that the next error is caught earlier.\nFigure 2. Even with the same model, results change with how goals, context, tools, validation, rubric improvement, and records are connected.\nResearch by METR reported a trend in which the length of well-defined software, machine-learning, and cybersecurity tasks that AI can complete with a 50% success rate doubled roughly every seven months.[2] METR also cautions that estimates beyond 16 hours are difficult to trust with its current task suite. Because the measured tasks are technical work with clear success criteria, this trend must not be converted directly into a percentage of jobs automated. Even so, the direction matters. As AI takes on longer tasks, the human role shifts toward setting goals and boundaries, designing intermediate checkpoints, and reversing failures. If organizational AI use remains confined to one-off questions and answers through the end of 2027, this transition will have been slower than expected.\n3. Bundles of Tasks and Entry-Level Pathways Will Change Before Jobs Do # An occupation consists of many kinds of work, and AI will not take over every task at the same speed. A 2025 analysis by the International Labour Organization (ILO) and Poland’s NASK estimated that one in four workers worldwide was in an occupation exposed to generative AI, while concluding that tasks within jobs were more likely to be transformed than entire jobs replaced.[3]\nWhat concerns me in particular is the entry pathway. Until now, beginners have developed expert judgment by organizing materials, writing first drafts, and doing basic coding. If AI begins by taking over precisely those tasks, beginners may produce results more quickly while skipping the process through which judgment develops. The challenge for education and organizations in 2027 is likely to be not “whether to ban AI,” but how to build fundamentals and the muscles of judgment while still using AI. This is my forecast, not an observed fact. I should revise it if the work and evaluation of beginners remain largely unchanged in 2027.\n4. Scarcity Will Shift from Producing Answers to Setting Goals, Verifying, and Taking Responsibility # As the cost of producing an answer falls, four things become scarce: the ability to decide what problem to solve, to read the context of the field, to question and verify a plausible result, and to put one’s name behind the final decision and take responsibility for it. AI can assist with these too, but assistance is not delegation.\nIn a study by Microsoft Research collaborators covering 936 cases of AI use reported by 319 knowledge workers, greater confidence in AI was associated with lower reported critical-thinking effort, while the center of cognitive effort shifted from generation to verification, integration, and task management.[4] Because this was a self-report survey, it does not establish causality. It does, however, reveal a design risk: the more people trust AI, the looser their review may become.\nWhether AI weakens or deepens thinking depends less on AI itself than on the harness we build around it. A system that encourages people to submit the first answer outsources thinking. One that requires competing hypotheses and traceable evidence expands it.\n5. We Need to Prepare Domain Knowledge, Questions, Harnesses, Verification, and Trust # Memorizing the menus of a particular model is not durable preparation. We need to build five assets that remain valuable even as tools change.\nDomain knowledge — An internal map for judging whether an AI answer falls within a sensible range. You should be able to explain the core concepts, how the data are produced, and which errors occur frequently. Questions and success criteria — Write them down before delegating the task. What decision is this work intended to support? What is the cost of being wrong? What evidence is required before the task is complete? Without criteria, the most fluent answer can look like the best one. A personal harness and rubric — Take one task you repeat each week and attach a goal, materials, sequence, checkpoints, and records. Create a first rubric covering factual accuracy, links to primary sources, counterevidence, and reproducibility. When a failure is discovered, revise the rubric as well as the prompt. A verification routine — Check numbers by hand on a small sample, confirm important facts in primary sources, and record competing hypotheses and failure conditions before making a decision. A second answer from the same model is not independent verification. Trust and responsibility — Record who approved the result and what remains unknown. Do not use a result you cannot explain in an important decision. Figure 3. The most durable preparation is not a particular AI tool, but domain knowledge, questions, a thinking harness, validation, and trust.\nThis month, choose just one recurring task. Perform it once without AI and record the time and errors. Create your first rubric. Then perform it with AI, adding two verification points, and run it again after adding any missed failure conditions to the rubric. The key is to compare not only speed, but whether both the error rate and your own understanding improve.\nThe forecast in this article should be evaluated in the same way. If agents have not become more reliable on unstructured tasks by the end of 2027 and entry-level learning pathways have barely changed, I should revise this forecast.\nConclusion — Do Not Hand Your Thinking to AI; Build a Structure That Deepens It # Do not try to become someone who produces answers faster than AI. Become someone who works with AI to ask better questions, verify more rigorously, and take responsibility for the result.\nUsing AI will not be remarkable in itself in 2027. The difference will lie in what was not delegated, where a person intervened, and what evidence and evaluation criteria were recorded.\nWhat recurring task in your work would you turn into a small harness first?\nRelated Articles # If AI Writes All the Code, Do We Still Need to Learn Python? — From Writing to Verification Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals Sources # [1] Stanford Institute for Human-Centered Artificial Intelligence. The 2026 AI Index Report. https://hai.stanford.edu/ai-index/2026-ai-index-report\n[2] METR. Task-Completion Time Horizons of Frontier AI Models (updated May 8, 2026); Kwa, T. et al. Measuring AI Ability to Complete Long Tasks. https://metr.org/time-horizons/ · https://arxiv.org/abs/2503.14499\n[3] International Labour Organization \u0026amp; NASK. Generative AI and Jobs: A Refined Global Index of Occupational Exposure (2025). https://www.ilo.org/publications/generative-ai-and-jobs-refined-global-index-occupational-exposure\n[4] Lee, H.-P. H. et al. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI \u0026lsquo;25. https://doi.org/10.1145/3706598.3713778\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"July 17, 2026","externalUrl":null,"permalink":"/en/posts/2026-07-17-ai-era-2027/","section":"Writing","summary":"In 2027, what may matter more than whether we use AI is the ability to design goals, context, tools, and verification as one system—and to improve the rubrics used to evaluate the results.","title":"What Will Change First in the AI Era of 2027? — How Work Will Change Faster Than Jobs","type":"posts"},{"content":" We Have the Code, but No One Who Understands It # These days, I often encounter the same scene while evaluating term projects at school or working on research projects.\nA student brings in analysis code. It runs without errors, and the resulting graphs look plausible. But when I ask, “Why did you handle this part this way?” the student cannot answer. AI wrote the code, and often even the person who brought it in does not know exactly what it does.\nWe have the code, but no one who understands it. So no verification takes place either.\nWhenever I encounter this scene, one question comes to mind. It is also one of the questions I hear most often today from students and professionals deciding whether to learn to code:\n“If AI is going to write all the code anyway, do I still need to learn Python?”\nToday, I will answer that question directly.\nThe Claim Is Half Right # Let us first acknowledge what is true. Code-generating AI is good. Even without knowing the syntax, you can describe what you want in ordinary language and receive code that runs. Tasks that once took half a day can indeed take only a few minutes.\nI am no exception. I use code-generating AI in development nearly every day. We are already past the stage of asking whether it is acceptable to use it. For me, AI coding has become not an option but an essential tool.\nSo I do not intend to shut down the question by saying, “The fundamentals are still important.” That statement is not wrong, but it explains nothing. Instead, let us change the question:\nNot “Can I write code?” but “When should I trust the result this code produces, and when should I question it?”\nIt is the same question this blog returns to repeatedly. This is precisely why Python still matters even when AI writes the code.\nThe Problem Is Not Writing Code, but Trusting It # Errors in AI-generated code come in two forms.\nFirst, errors that fail loudly. The program raises an error and stops. This is actually a good kind of error: the code itself tells us that something is wrong. Feed the error message back to the AI and it will fix the problem in most cases. This is a stage that people can often get through somehow even without much Python knowledge.\nSecond, errors that fail silently. The program runs all the way through without an error and produces plausible numbers and graphs. The only problem is that the numbers are wrong. Because the program never stops, no one notices.\nThe second kind is the real problem, and it becomes more dangerous as AI code generation spreads. Think back to the opening scene. If no one “wrote” the code, there may be no one prepared to “question” it either.\nFigure 1. Loud errors reveal themselves; silent errors may never be discovered.\nCode That Runs Is Not Necessarily Code That Is Correct # Here are several common examples from data analysis of what it means for code to fail silently. Every one of them can run without an error.\nMissing values disappear without a sound. When you calculate an average with Python’s pandas library, empty values (NaN) are excluded by default. An average for a day with half its observations missing looks just as valid as an average for a complete day. Mismatched units do not stop the calculation. Even if one file uses millimeters and another uses centimeters, Python does not complain. It simply multiplies and adds the values. Merging data can multiply the number of rows. If you join two tables on the wrong key, the same data may be duplicated several times, yet the code still completes successfully. With more apparent observations, the statistics become more “confidently” wrong. Time can shift by a day. Anyone who works with data is likely to have seen every date move silently because of one timezone-handling choice. These errors do not occur because AI is uniquely incapable. Humans make them too. The difference is that when a person writes the code, at least someone made the choice. When AI writes it, the choice can be made without anyone noticing. Whether to drop missing values, fill them in, or discard the entire day is not a coding question but an analytical judgment—and that judgment becomes hidden inside a default setting.\nFigure 2. Even when code does not stop, its result can still be quietly wrong.\nIt Will Invent Data If That Is What It Takes to Produce an Answer # There is one form of silent error that I consider especially dangerous.\nToday’s AI is best understood as a tool trained with a strong tendency to produce an answer somehow. Recent models may ask follow-up questions or stop, but when blocked they still have a tendency to invent something plausible. This is commonly called hallucination, and it does not occur only in prose. It happens in code as well.\nHere is a pattern I have witnessed several times in real projects. I asked AI to write data-analysis code, but when required data were missing or the format did not match, instead of stopping and asking a question, it filled the gaps with plausible values—or even created data that did not exist—to complete the result. On the surface, the analysis looked successful. Without verification, data that never existed in the real world remained embedded in the result.\nI build AI systems that predict infectious diseases and environmental conditions. In these fields, “silently invented data” does more than make a graph look slightly odd. Predictions become evidence used in real-world judgments.\nThis risk applies equally to us. As I said earlier, my team and I also use code-generating AI as an essential tool. The danger does not come from code being written by AI; it comes from code not being verified. By that standard, the code we produce is no exception. That is why the more plausible a result looks, the more deliberately I try to question it one more time.\nFigure 3. When AI fills a gap with a plausible value, the fabrication can be hard to see.\nSo How Much Do We Need to Learn? # If your reaction so far is, “So the answer is that I need to study coding hard after all,” you are only half right. My answer is slightly different: the center of gravity of “the Python we need to learn” has shifted in the age of AI.\nIn the past, the focus was the ability to fill a blank screen with code—the ability to write. That has become the part AI is best at replacing. What remains, however, has not been replaced:\nA feel for the shape of data. Tables, lists, dictionaries—you need to be able to picture the form of your data in your mind before you can suspect what may be missing and where. The ability to read code written by someone else. AI-generated code is also “code written by someone else.” You do not need to understand every line. You need to read well enough to locate the points worth questioning: “What happens to missing values here?” “Is this the right join key?” A habit of testing on a small scale. Before running the full dataset, run ten rows and compare the output with a value calculated by hand. This is closer to an attitude than a knowledge of syntax, but it cannot be practiced without at least some Python. The ability to reproduce the same result. A result that cannot be reproduced cannot be verified. What these four abilities share is not “writing,” but reading and questioning. The amount we need to learn may even have decreased. The direction has changed.\nFigure 4. In the age of AI coding, Python is increasingly a tool for reading and questioning, not just writing.\nWhen to Question AI-Generated Code # To summarize, here are five warning signs I look for—questions to ask yourself when you see the output of AI-generated analysis code.\nThe result is too good. If accuracy suddenly jumps, first check for data leakage—information about the correct answer accidentally included in the inputs—before celebrating an improvement in the model. No one checked the row counts. If you did not count the rows before and after joins or filters, the result is not yet ready to be trusted. No one knows what happened to missing values. If you cannot say how many values were missing or how they were handled, averages and trends may have been silently distorted. Check especially whether the AI “filled in” the gaps. Not a single value was checked by hand. A result that has not been compared with a hand calculation on even one tiny subset has run, but it has not been verified. No one can explain why it was handled that way. Pick any choice in the code and ask “Why?” If you cannot answer, you do not yet own the analysis. Conclusion — Back to the Original Question # “If AI writes everything, do I still need to learn Python?”\nMy answer is this: we will entrust more and more of the work of writing code to AI. The more we do so, the more the work of questioning and verifying that code will belong to people. The meaning of learning Python is shifting from the skill of writing code to the ability to decide whether to trust it.\nIn the previous article, I discussed why AI predictions should not answer with “one number.” When to trust an AI answer and when to question it is not a concern limited to predictive models. It begins with a single line of code written by AI.\nWhat about you? How much do you check before using code or analytical results produced by AI? Tell me in the comments, and I will draw on your responses in the next article.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know.\n","date":"July 14, 2026","externalUrl":null,"permalink":"/en/posts/2026-07-14-python-still-needed/","section":"Writing","summary":"In an age when AI can write all the code, the reason to learn Python lies not in writing code but in verifying it. When should we trust or question code that fails silently, or AI that invents missing data to produce an answer?","title":"If AI Writes All the Code, Do We Still Need to Learn Python? — From Writing to Verification","type":"posts"},{"content":" When You Do Not Know What You Do Not Know # Do you know the most difficult moment when you are studying?\nIt is not when you get something wrong. It is when you do not even know what you do not know.\nA wrong answer is almost easier. At least you know what to review. What is truly frustrating is feeling as though you know everything but being unable to write an answer—and then being unable to explain what you do not understand when someone asks.\nI have found that learning begins to improve at a particular point: the moment you can say that you do not know something. From then on, you know what to look up and whom to ask.\nMost AI systems today are still at the stage before that. They produce a single answer without knowing what they do not know.\nI research predictive models for infectious diseases (avian influenza) and environmental problems (algal blooms and particulate matter), and I have applied them in real-world settings. Along the way, I learned one thing: what people in the field really want to know is not the prediction itself, but how much they can trust it.\nWhen I bring a prediction into the field, there is something people ask about before the number itself: what evidence supports this value? And what I can usually offer in response is past performance metrics.\nThis is the first article in the “Trustworthy AI” series. One promise runs through the entire series: to make AI say, “I don’t know.” As a starting point, I will discuss the difference between point and interval estimates.\nTwo Answers to “What Time Will You Arrive?” # Let us begin with the terminology in plain language.\nA friend asks, “What time will you arrive?”\nThere are two possible answers. One is “Exactly 3:00.” The other is “Sometime between 3:00 and 3:30, almost certainly.”\nThe first is what statistics calls a point estimate: selecting one value as the most plausible answer. The second is an interval estimate: giving a range as the answer, together with how much confidence to place in that range.\nIn daily life, we already speak in intervals because traffic might delay us. Yet, strangely, most AI systems still answer only with points. “The algal concentration at this site tomorrow will be 42.” That is all. They do not tell us how much to trust that value.\nThe Illusion Created by a Single Number — Point Estimates # Point estimates are appealing because they are clear. “Tomorrow’s concentration: 42” looks tidy on a presentation slide.\nBut there is a trap. When predicting a continuous value, the probability that the prediction will be exactly correct is effectively close to zero. The actual value might be 41.3 or 48.9. The real questions are therefore these:\nWhere around 42 is the actual value likely to fall? How wide is that surrounding range? Even with the same point estimate of 42, “it is likely to be between 40 and 44” tells a completely different story from “it could be anywhere between 20 and 70.” The first 42 can support action; the second is merely a reference point. A single point does not distinguish between them. The clearer the number appears, the stronger the illusion of certainty becomes.\n“It Was 92% Accurate Last Year” Is Not Persuasive Enough # When we argue that a model is useful in practice, we really have only one card to play: past performance metrics. We tested it on data from the past several years, and this is how often it was right.\nBut the person sitting across the table is really asking something else:\n“So how much do you trust this number, at this site, today?”\nPast performance is a statement averaged over the whole. It tells us how a model performed on average across thousands of predictions. What the person is looking at, however, is one case, here and now.\nModels have easy days and hard days. Some sites have dense observations; others have sparse ones. Some conditions are familiar; others have never been seen before. An average accuracy of 92% does not mean that today’s prediction is 92% certain.\nA point estimate cannot bridge this gap. It produces one number and adds, “Our model performs well on average.” The person responsible for acting on it still lacks the evidence needed to decide whether it is safe to rely on that statement.\nA prediction interval provides that evidence. It is not a statement about past performance as a whole, but a statement about this input now. If today’s data are poor, the interval widens. If the conditions are familiar, it narrows. It is the model saying, “I am not very confident today either.”\nCan we trust the interval itself? That is a good question—and the subject of the next article.\nStatistics Has Already Gone Through This Transition # Statistics passed through this problem a century ago.\nIn the early twentieth century, statisticians focused on refining theories for finding the single best value: point estimation. The person who changed that trajectory was the Polish-born statistician Jerzy Neyman. In a 1934 paper, he introduced the concept of the confidence interval, and in a 1937 paper he developed it into a systematic theory.[1][2]\nThat work established a framework for reporting not “one answer,” but “a range likely to contain the answer, together with how much the procedure can be trusted.”\nConfidence intervals subsequently became a standard language of science. Today, research papers and clinical-trial results report intervals alongside individual values. The most familiar example is an opinion poll. Next to “candidate support: 45%” we expect wording such as “margin of error: ±3 percentage points at the 95% confidence level.” Society has, in effect, agreed that a poll reported without this information is incomplete.\nWe moved from stating one number to stating both the number and the degree of confidence we can place in it. Keep this transition in statistics in mind. The same thing is now happening in AI.\nAI Is in the Middle of the Same Transition # The default output of today’s AI, especially deep-learning models, is a point estimate: tomorrow’s concentration, next week’s outbreak risk, or the interpretation of an image. Most systems return a single value or label.\nThat is why uncertainty quantification has become an important area of research. Its goal is to have a model produce not only an answer, but also an indication of how confident or uncertain it is about that answer. Comprehensive review articles have surveyed this field; two such reviews have been cited more than 2,000 and 1,000 times, respectively, as of July 2026.[3][4] That is evidence that this is an active area of research.\nIn my own words: just as statistics moved from point estimates to confidence intervals, AI is moving from an age of point estimates to an age of intervals. Prediction is shifting from “tomorrow’s algal concentration will be 42” to “42, within this range, at this level of confidence.”\nThis Transition Will Not Reverse — Because the Areas of Application Have Changed # Why do I believe this is more than a passing trend? Because the areas in which AI is being applied have changed.\nUntil recently, AI was used mainly in convenience-oriented areas such as recommendations, advertising, and search. The cost of an incorrect prediction in these settings is small. If a film recommendation does not match your taste, you simply do not watch it. There was little need to say, “I am 87% confident in this recommendation.”\nNow AI is entering medicine, infectious-disease control, environmental management, and infrastructure—areas where mistakes carry serious costs. Water-treatment operations, the allocation of disease-control resources, diagnostic support: if predictions inform decisions such as these, “how confident are you?” is not optional metadata. It is an essential part of the output. The greater the cost of being wrong, the less usable a prediction without uncertainty information becomes.\nRegulation points in the same direction.\nIn Annex III, the European Union’s AI Act (Regulation (EU) 2024/1689) classifies AI used as a safety component in managing and operating the supply of water, gas, heating, or electricity as “high-risk.”[5] The law has begun to treat AI used for convenience differently from AI used where safety is at stake.\nReduced to one sentence, the legal direction is this: the more important the setting in which AI is used, the more clearly it must explain how much its judgment can be trusted.\n“So What Can We Do with It?” # When we say we have built a model to predict something, one question almost always comes back from the field. I hear it often as well:\n“So what can we do with that prediction?”\nA prediction is not an action. It becomes valuable only when it leads to a decision about what to change—what is often called an intervention. What determines that decision is the width of the uncertainty interval.\nIf the interval is narrow, we can act directly on the prediction. If the interval is wide, we may collect more observations or choose a conservative response that prepares for the worst case. Even with the same prediction of 42, what we do next changes depending on whether the interval is 40–44 or 20–70. In the first case, we proceed as planned. In the second, we measure again or leave more room for error.\nAn interval is therefore not merely a display of humility by the model; it is guidance for the user’s next action. A model that offers only a point estimate leaves this entire judgment to the user’s intuition. It predicts, but does not help decide what to do—a partial output at best.\nFigure 1. The point estimate is identical, but the interval width changes the next action.\nDoes High Uncertainty Mean a Bad Model? # At this point, a reasonable objection may arise: if uncertainty is high, does that not simply mean the model is poor?\nI do not see it that way.\nReturn to the example of studying. One student knows exactly what they do not understand. Another thinks they know everything but repeatedly gets answers wrong. Which is the better student?\nA model that reports high uncertainty is not necessarily a bad model; it is a model that knows its limits. The truly dangerous model is one that is wrong with confidence. The first allows us to prepare; the second leaves us defenseless.\nOf course, an interval that is too wide makes a prediction less useful in practice. That is true. Even then, however, we know that the prediction cannot support a decision. That is a very different position from being unaware of what we do not know.\nConclusion — The Promise of This Series # In one sentence: good AI does not merely give accurate answers; it also tells us how likely it is to be accurate.\nLearning works the same way. It begins when we can say that we do not know what we do not know. AI is now crossing that threshold.\nIt took statistics a generation to move from point estimates to confidence intervals. AI is in the middle of that transition now, and the shift will accelerate as AI expands into safety-critical areas. In this series, I will explain, one step at a time, how to make AI say, “I don’t know.”\nThe next article will be about calibration. When a model says it is “90% confident,” is that 90% really 90%? We will begin with that question.\nWhat about your own field? When you receive an AI prediction, have you ever wondered how much you should trust it? Tell me in the comments, and I will draw on your responses in the next article.\nSources # [1] Neyman, J. (1934). On the Two Different Aspects of the Representative Method: The Method of Stratified Sampling and the Method of Purposive Selection. Journal of the Royal Statistical Society, 97(4), 558–625. doi:10.2307/2342192 — The paper in which the concept of confidence intervals was first introduced.\n[2] Neyman, J. (1937). Outline of a Theory of Statistical Estimation Based on the Classical Theory of Probability. Philosophical Transactions of the Royal Society A, 236(767), 333–380. doi:10.1098/rsta.1937.0005 — The paper that developed confidence intervals into a systematic theory.\n[3] Abdar, M. et al. (2021). A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76, 243–297. doi:10.1016/j.inffus.2021.05.008\n[4] Gawlikowski, J. et al. (2023). A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56, 1513–1589. doi:10.1007/s10462-023-10562-9\n[5] Regulation (EU) 2024/1689 (EU AI Act), Article 6 and Annex III. Annex III, point 2: critical infrastructure—AI systems intended to be used as safety components in the management and operation of the supply of water, gas, heating, or electricity. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng\nDisclosure of Interests and Responsibility # Predictions always contain uncertainty. They must not be used as the sole basis for decisions on disease-control or environmental policy.\nDisclosure of Interests and Responsibility\nDonghyun Lee is a professor in the Division of Social Science \u0026amp; AI at Hankuk University of Foreign Studies and the CEO of AI Korea Inc. The views expressed in this article are the author’s own and do not represent the official position of his affiliated institutions or the organizations commissioning his research projects.\nThis article is intended for general informational and educational purposes. It is not advice or a policy recommendation for any particular matter, and must not be used as the sole basis for real-world decisions.\nIf you find a factual error, please let me know. I will correct the original and disclose the correction.\n","date":"July 12, 2026","externalUrl":null,"permalink":"/en/posts/2026-07-07-ai-uncertainty-interval/","section":"Writing","summary":"Just as statistics moved from point estimates to confidence intervals, AI is entering an era in which it must report not just one answer, but how much that answer can be trusted. Part 1 of the Trustworthy AI series.","title":"Why AI Predictions Should Not Give a Single Number — From Point Estimates to Confidence Intervals","type":"posts"},{"content":"When AI says, “Avian influenza will spread next week,”\nhow far should we trust that answer?\nI study that “how far.” My work centers on trustworthy AI: when predictions and simulations from AI can be trusted in high-stakes physical and social systems—and when they should be rejected. I have used deep learning to forecast problems such as algal blooms and infectious disease, where errors can be costly. More recently, I have been testing whether LLM-based synthetic agents that respond like people are valid tools for social and behavioral research.\nOne accuracy score is not enough to make AI trustworthy. To use AI in practice, we need to see both the basis for its answer and the conditions under which it breaks down. The writing, teaching, and research on this site all return to that question.\nI am currently developing this perspective into two Korean-language books in the AI ERA SERIES: Python Fundamentals in the Age of AI and Python Data Visualization in the Age of AI, scheduled for publication in September 2026. About the books →\nWhat I do # Education \u0026amp; Research\nProfessor, Hankuk University of Foreign Studies I study and teach at the intersection of AI and society in the Division of Social Science \u0026amp; AI Convergence. I earned my bachelor’s, master’s, and doctoral degrees in engineering from KAIST.\nReal-World Application\nCEO, AI Korea Inc. I founded an AI startup and directly develop and operate practical AI services in infectious disease and the environment.\nResearch focus # Explainable AI (XAI) Uncertainty quantification Infectious-disease forecasting Environmental forecasting Deep learning Validity of LLM-based synthetic agents I have participated in projects that tackle real-world forecasting problems, including the national Digital Columbus project, and have led numerous research projects as principal investigator.\nBackground # Previously, I was an associate professor at Tech University of Korea and an associate research fellow on the Big Data Research Team at the Korea Environment Institute. I also served on the Environmental Economy Subcommittee of the Ministry of Environment’s Central Environmental Policy Committee. I have taught programming and data analysis at universities for ten years, and I draw on experience from both the classroom and the startup field while writing the AI ERA SERIES.\nPublications and awards # I have published more than twenty research papers in international journals, including Journal of Cleaner Production, Technological Forecasting and Social Change, and Expert Systems with Applications. I received the President’s Award from the Korea Environment Institute for developing a multimodal-LLM conversational AI specialized for environmental research. In 2008, I received the Talent Award of Korea, a presidential award. I also serve on the editorial board of Journal of Innovation \u0026amp; Knowledge.\nContact For speaking, advisory, or research collaboration inquiries, email donghyunlee.ai@gmail.com.\n","date":"July 11, 2026","externalUrl":null,"permalink":"/en/about/","section":"Donghyun Lee’s Trustworthy AI Notes","summary":"Donghyun Lee — professor at Hankuk University of Foreign Studies and founder of AI Korea Inc., applying trustworthy AI to high-stakes forecasting.","title":"About","type":"page"}]