CodalSearch this book — or all of Codal…⌘K
nydus/The Logic of Chance, 3rd EditionPublic

This work examines the physical foundations of probability by analyzing the formation and behavior of statistical series. It explores the nature of laws of error, the processes of causation, and the empirical methods required to establish and prove probabilistic data.

Page 292 of 310
Table of Contents

THE THEORY OF THE AVERAGE AS A MEANS OF APPROXIMATION TO THE TRUTH.

in the end, since positive and negative errors are supposed to be equal and opposite,—will itself be an ‘error’, every magnitude of which will have a certain assignable probability or facility of occurrence. What we do is to assign the modulus of these errors. The actual result again is simple. If c had been the modulus of the single errors, that of the sum or difference of the averages of m and n of them will be

§ 20. So far, the problem under investigation has been of a direct kind. We have supposed that the ultimate mean value or central position has been given to us; either à priori (as in many games of chance), or from more immediate physical considerations (as in aiming at a mark), or from extensive statistics (as in tables of human stature). In all such cases therefore the main desideratum is already taken for granted, and it may reasonably be asked what remains to be done. The answers are various. For one thing we may want to estimate the value of an average of many when compared with an average of a few. Suppose that one man has collected statistics including 1000 instances, and another has collected 4000 similar instances. Common sense can recognize that the latter are better than the former; but it has no idea how much better they are. Here, as elsewhere, quantitative precision is the privilege of science. The answer we receive from this quarter is that, in the long run, the modulus,—and with this the probable error, the mean error, and the error of mean square, which all vary in proportion,—diminishes

inversely as the square root of the number of measurements or observations. (This follows from the second of the above formulæ.) Accordingly the probable error of the more extensive statistics here is one half that of the less extensive. Take another instance. Observation shows that “the mean height of 2,315 criminals differs from the mean height of 8,585 members of the general adult population by about two inches” (v.[ TN: space] Edgeworth, Methods of Statistics: *Stat.[ TN: space] Soc.[ TN: space] Journ.*[ TN: space] 1885). As before, common sense would feel little doubt that such a difference was significant, but it could give no numerical estimate of the significance. Appealing to science, we see that this is an illustration of the third of the above formulæ. What we really want to know is the odds against the averages of two large batches differing by an assigned amount: in this case by an amount equalling twenty-five times the modulus of the variable quantity. The odds against this are many billions to one.

§ 21. The number of direct problems which will thus admit of solution is very great, but we must confine ourselves here to the main inverse problem to which the foregoing discussion is a preliminary. It is this. Given a few only of one of these groups of measurements or observations; what can we do with these, in the way of determining that mean about which they would ultimately be found to cluster? Given a large number of them, they would betray the position of their ultimate centre with constantly increasing certainty: but we are now supposing that there are only a few of them at hand, say half a dozen, and that we have no power at present to add to the number.

In other words,—expressing ourselves by the aid of graphical illustration, which is perhaps the best method for the novice and for the logical student,—in the direct problem we merely have to draw the curve of frequency from a knowledge of its determining elements; viz.[** TN: space] the position of the centre, and the numerical value of the modulus. In the inverse problem, on the other hand, we have three elements at least, to determine. For not only must we, (1), as before, determine whereabouts the centre may be assumed to lie; and (2), as before, determine the value of the modulus or degree of dispersion about this centre. This does not complete our knowledge. Since neither of these two elements is

292