THE

INTERNATIONAL SERIES

OF

MONOGRAPHS ON PHYSICS

GENEBAL EDITORS

fR. H. FOWLER, P. KAPITZA N. F. MOTT, E. C. BULLARD

THE INTERNATIONAL SERIES OF MONOGRAPHS ON PHYSICS

GENERAL EDITORS The late Sib RALPH FOWLER N. F. MOTT

Henry Overton Wills Professor of Theoretical Physics in the University of Bristol.

Already Published

THE THEORY OF ELECTRIC AND MAGNETIC SUSCEPTIBILITIES.

By J. H. VAN VLECK. 1932.

THE THEORY OF ATOMIC COLLISIONS. By n. p. mott and h. s. w. MASSEY. Second edition. 1949.

RELATIVITY, THERMODYNAMICS, AND COSMOLOGY. By r. c. tol- MAN. 1934.

CHEMICAL KINETICS AND CHAIN REACTIONS By n. sembnoff. 1935. RELATIVITY, GRAVITATION, AND WORLD-STRUCTURE. By e. a. MILNE. 1935.

KINEMATIC RELATIVITY. A sequel to Relativity, Oravitation, and World- Structure. By e. a. milne. 1948.

THE QUANTUM THEORY OF RADIATION. By w. heitleb. Second edition. 1945.

THEORETICAL ASTROPHYSICS: ATOMIC THEORY AND THE ANALYSIS OP STELLAR ATMOSPHERES AND ENVELOPES. By 8. BOSSELAND. 1936.

ECLIPSES OF THE SUN AND MOON. By sir frank dyson and b. v. d. R. WOOLLEY. 1937.

THE PRINCIPLES OF STATISTICAL MECHANICS. By r. c. tolman. 1938. ELECTRONIC PROCESSES IN IONIC CRYSTALS. By n. f. mott and R. w. GURNEY. Second edition. 1948.

GEOMAGNETISM. By s. chapman and j. bartels. 1940. 2 vols.

THE SEPARATION OF GASES. By m. ruhemann. Second edition. 1949. KINETIC THEORY OF LIQUIDS. By j. prenkkl. 1946.

THE PRINCIPLES OF QUANTUM MECHANICS. By p. a. m. dirac. Third edition. 1947.

THEORY OF ATOMIC NUCLEUS AND NUCLEAR ENERGY SOURCES. By 0. OAMOw and c. L. cbitchfield. 1949. Being the third edition of

STRUCTURE OF ATOMIC NUCLEUS AND NUCLEAR TRANSFORMATIONS.

THE PULSATION THEORY OF VARIABLE STARS. By s. bosseland. 1948.

COSMIC RAYS. By l. jXnossy. New edition in preparation.

THEORY OF PROBABILITY. By Harold Jeffreys. Second edition. 1948.

P. KAPITZA E. C. BULLARD Director of the National Physical Laboratory, Teddington.

THEORY OF

PROBABILITY

BY

HAROLD JEFFREYS

M.A., D.Sc., F.R.^.

PLUMIA?f PROFESSOR OF ASTRONOMY UNIVERSITY OF CAMBRIDGE

SECOND EDITION

OXFORD

AT THE CLARENDON PRESS

Oxford University Press ^ Amen Hovse, London E.C. 4

GLASGOW NEW YORK TORONTO MELBOURNE WELLINGTON BOMBAY CALCUTTA MADRAS CAPETOWN

Geoffrey Cumherlege, Publisher to the University

FIRST EDITION 1939 SECOND EDITION I948

Reprinted lithographically in Great Britain at the UNIVERSITY PRESS, OXFORD, I950 from sheets of the second edition

PREFACE TO THE SECOND EDITION

In the circumstances that have prevailed in the world since the appear- ance of this book, it is a welcome indication of increasing interest in the principles of scientific method that a second edition has been required. I have taken the opportunity to add some arguments that go far towards establishing the consistency of the product rule and therefore of the principle of inverse probability. A theory of invariance has been developed and applied to problems of estimation and significance, thus establishing the possibility of a consistent rule for stating prior proba- bilities over large parts of the subject. I am not satisfied that it is the only such rule or even the best one, but think that enough progress has been made to indicate that the attempt is worth pursuing.

I have not attempted to answer explicitly the criticisms made by reviewers, because on examination I found that they were all dealt with in the book already. What does strike me as remarkable is that no mention was made of the fact that the book jpontained useful methods of treatment of several problems of practical impc^rtance. I have still not gathered what distinction those statisticians who do not accept the epistemological approach draw between estimation problems and significance tests, or whether they think that they are saying anything about a hypothesis when they reject it. So far as I can judge from their pronouncements, they provide themselves with no reason against continuing to make predictions from it.

Several recent writers, especially in the United States, have described me as a follower of the late Lord Ke3mes. Without wishing to disparage Keynes, I must point out that the first two papers by Wrinch and me in the Philosophical Magazine of 1919 and 1921 preceded the publication of Keynes’s book. What resemblance there is between the present theory and that of Keynes is due to the fact that Broad, Keynes, and my col- laborator had all attended the lectures of W. E. Johnson. Keynes’s distinctive contribution was the assumption that probabilities are only partially ordered ; this contradicts my Axiom 1. I gave reasons for not accepting it in Scientific Inference. Keynes himself withdrew it in his biographical essay on F. P. Ramsey.

I have to thank several correspondents for suggesting corrections, especially Dr. H. Chojnacki-Hanani. Mr. P. H. Diananda of Caius College and Mr. V. S. Huzurbazar of Fitzwilliam House, Cambridge, have helped greatly in the proof-correction. g j

ST. JOHN’S COLLEGE, CAMBBIDGE

October 1947

PREFACE TO THE FIRST EDITION

The chief object of this work is to provide a method of drawing infer- ences from observational data that will be self-consistent and can also be used in practice. Scientific method has grown up without mucli attention to logical foundations, and at present there is little relation between three main groups of workers. Philosojihers, mainly interested in logical principles but not much concerned with specific a])})lications, have mostly followed in the tradition of Bayes and Lajfiace ; but with the brilliant exception of Professor C. D. Broad have not paid much attention to the consequences of adhering to the tradition in detail. Modern statisticians have develojied extensive mathematical techniques, but for the most part have rejected the notion of the probability of a hypothesis, and thereby deprived themselves of any way of saying precisely what they mean when they decide betwe^en hyjiotheses. Physicists have been described, by an experimental physicist who has devoted much attention to the matter, as not only indifferent to funda- mental analysis but actively hostile to it ; and with few excejitions their statistical technique has hardly advanced beyond that of Lajilace. In opposition to the statistical school, they and some other scknitists are liable to say that a hypothesis is definitely pi’oved by observation, which is certainly a logical fallacy ; most statisticians appear to regard observations as a basis for possibly rejecting hypotheses, but in lio case for supporting them. The latter attitude, if adojked consistently, would reduce all inductive inference to guessw^ork ; the former, if adopted consistently, wamld make it impossible ever to alter the hypo- theses, however badly they agreed with new^ evidence. The Y^resent attitudes of most physicists and statisticians are diametrically oj)posed, but lack of a common meeting-ground has, to a very large extent, jjre- vented the opposition from being noticed. Nevertheless, both schools have made great scientific advances, in sjute of the fact that their fundamental notions, for one reason or the other, would make such advances impossible if they were consistently maintained.

In the present book I reject the attempt to reduce induction to deduction, which is characteristic of both schools, and maintain that the ordinary common-sense notion of probability is capable of j)recise and consistent treatment when once an adequate language is provided for it. It leads to the result that a precisely stated hypothesis may attain either a high or a negligible probability as a result of observa- tional data, and therefore to an attitude intermediate between those current in physics and statistics, but in accordance with ordinary

PREFACE TO THE FIRST EDITION vii

thought. Fundamentally the attitude is that of Bayes and Laplace, though it is found necessary to modify their hypotheses before some types of cases not considered by them can be treated, and some steps in the argument have been filled in. For instance, the rule for assessing probabilities given in the first few lines of Laplace’s book is Theorem 7, and the principle of inverse probability is Theorem 10. There is, on the whole, a ver}^ good agreement with the recommendations made in statistical practice ; my objection to current statistical theory is not so much to the way it is used as to the fact that it limits its scope at the outset in such a way that it cannot state the questions asked, or the answers to them, within the language that it ])rovides for itself, and must either appeal to a feature of ordinary language that it has declared to be meaningless, or else })roduce arguments within its own language that will not bear inspection.

The most beneficial result that I can hope for as a consequence of this work is that more attention will be paid to the precise statement of the alternatives involved in the questions asked. It is sometimes considered a paradox that the answer depends not onlj^ on the observa- tions but on the question ; it should be a platitude.

The theory is applied to most of the main problems of statistics, and a number of specific ap])lications are given. It is a necessary condition for their inclusion that they shall have interested me. As my object is to })roduce a general method I have taken examples from a number of subjects, though naturally there are more from physics than from biology and more from geophysics than from atomic ph3^sic8. It was, as a matter of fact, mostly with a view to geophysical applications that the theory was developed. It is not easy, however, to produce a statistical method that has application to only one subject; though intraclass correlation, for instance, which is a matter of valuable posi- tive discovery in biology, is usually an unmitigated nuisance in physics. It may be felt that many of the applications suggest further questions. That is inevitable. It is usually only when one group of questions has been answered that a further group can be stated in an answerable form at all.

I must offer my warmest thanks to Professor R. A. Fisher and Dr. J. Wishart for their kindness in answering numerous questions from a not very docile pupil, and to Mr. R. B. Braithwaite, who looked over the manuscript and suggested a number of improvements; also to the Clarendon Press for their extreme courtesy at all stages.

H. J.

ST. John’s college, cambbidge

CONTENTS

I. FUNDAMENTAL NOTIONS 1

II. DIRECT PROBABILITIES . . . . .47

III. ESTIMATION PROBLEMS . . . . .09

IV. APPROXIMATE METHODS AND SIMPLIFICATIONS . 168

V. SIGNIFICANCE TESTS: ONE NEW PARAMETER . . 220

VI. SIGNIFICANCE TESTS: VARIOUS COMPLICATIONS . 305

VII. FREQUENCY DEFINITIONS AND DIRECT METHODS . 341

VIII. GENERAL QUESTIONS . . . . .372

APPENDIX. TABLES OF X . . . .396

NOTE ON THE CONSISTENCY OF THE PRODUCT RULE . 405

NOTE ON THE INFINITE REGRESS ARGUMENT . . 407

INDEX ........ 408

I

FUNDAMENTAL NOTIONS

They say that Understanding ought to work tlio rules of* right reason. These rules are, or ought to bo, contained in Logie ; but the actual science of logic is conversant at present only with things either certain, iinpo.ssible, or entirely doubtful, none of which (fortunately) we havx^ to reason on. I'herofore the true logic for this world is the calculus of Probabilities, which t akes account of the magnitude of the probability which is, or ouglit to be, in a reasonable man’s mind.

J. Clerk Maxwell

1.0. The fundamental problem of scientific progress, and a fundamental one of everyday life, is that of learning from experience. Knowledge obtained in this way is partly merely description of what we have already observed, but part consists of making inferences from past experience to predict future experience. This part may be called generalization or induction. It is the most important part; events that are merely described and have no apparent relation to others may as well be for- gotten, and in fact usually are. The theory of learning in general is the branch of logic known as epistemology. A few' illustrations w'ill indicate the scope of induction. A botanist is confident that the plant that grows from a mustard seed w ill have yellow' flow ers with four long and two short stamens, and four petals and sepals, and this is inferred from previous instances. The Naiitical Abnanac's predictions of the positions of the planets, an engineer’s estimate of the output of a new dynamo, and an agricultural statistician’s advice to a farmer about the utility of a fertilizer are all inferences from past experience. When a musical composer scores a bar he is expecting a definite series of sounds when an orchestra carries out his instructions. In every case the inference rests on past experience that certain relations have been found to hold; and those relations are then applied to new^ cases that were not part of the original data. The same applies to my expectations about the flavour of my next meal. The j)rocess is so habitual that we hardly notice it, and we can hardly exist for a minute without carry- ing it out. On the rare occasions when anybody mentions it, it is called common sense and left at that.

Now such inference is not covered by logic, as the w'ord is ordinarily understood. Traditional or deductive logic admits only three attitudes to any proposition: definite proof, disproof, or blank ignorance. But no number of previous instances of a rule will provide a deductive proof

3595.58

B

2 FUNDAMENTAL NOTIONS Chap. I

that the rule will hold in a new instance. There is always the formal possibility of an exception.

Deductive logic and its close associate, pure mathematics, have been developed to an enormous extent, and in a thoroughly systematic way indeed several ways. Scientific method, on the other hand, has grown up more or less haphazard, techniques being developed to deal with problems as they arose, without much attempt to unify them, except so far as most of the theoretical side involved the use of pure mathe- matics, the teaching of which required attention to the nature of some sort of proof. Unfortunately the mathematical proof is deductive, and induction in the scientific sense is simply unintelligible to the pure mathematician as such; in his unofficial capacity he may be able to do it very well. Consequently little attention has been paid to the nature of induction, and apart from actual mathematical technique the relation between science and mathematics has done little to develop a connected account of the characteristic scientific mode of reasoning. Many works exist claiming to give such an account, and there are some highly useful ones dealing with methods of treating observations that have been found useful in the past and may be found useful again. But when they try to deal with the underlying general theory they suffer from all the faults that modern pure mathematics has been try- ing to get rid of: self-contradictions, circular arguments, postulates used without being stated, and postulates stated without being used. Running through the whole is the tendency to claim that scientific method can be reduced in some way to deductive logic, which is the most fundamental fallacy of all: it can be done only by rejecting its chief feature, induction.

The principal field of application of deductive logic is pure mathe- matics, which pure mathematicians recognize quite frankly as dealing with the working out of the consequences of stated rules with no reference to whether there is anything in the world that satisfies those rules. Its propositions are of the form ‘If p is true, then q is true', irrespective of whether we can find any actual instance where p is true. The mathematical proposition is the whole proposition, ‘If p is true, then q is true’, which may be true even if p is in fact always false. In applied mathematics, as usually taught, general rules are asserted as applicable to the external world, and the consequences are developed logically by the technique of pure mathematics. If we inquire what reason there is to suppose the general rules true, the usual answer is simply that they are known from experience. However, this use of the

FUNDAMENTAL NOTIONS

3

§ 1.0

word ‘experience’ covers a confusion. The rules are inferred from past experience, and then applied to future experience, which is not the same thing. There is no guarantee whatever in deductive logic that a rule that has held in all previous instances will not break down in the next instance or in all future instances. Indeed there are an infinite number of rules that have held in all previous cases and cannot possibly all hold in future ones. For instance, consider a body falling freely under gravity. It would be asserted that the distance at time t below a fixed level is given by a formula of the type

6* = a+ut+\gt^. (1)

This might be asserted from observations of 5 at a series of instants That is, our previous experience asserts the proposition that a, u, and g exist such that

-- a+ut^+\gtl (2)

for all values of r from 1 to ii. But the law (1) is asserted for all values of t. But consider the law

where f{t) may be any function whatever that is not infinite at an}^ of

^1, and a, m, and g have the same values as in (1). There are an

infinite number of such functions. Every form of (3) will satisfy the set of relations (2), and therefore every one has held in all previous cases. But if we consider any other instant (which might be either within or outside the range of time between the first and last of the original observations) it will be possible to choose f{fn+i) such a way as to give s as found from (3) any value wdiatever at time Further, there will be an infinite number of forms of f(t) that would give the same value of /(f„4i), and there are an infinite number that would give different values. If we observe s at time we can choose to give agreement with it, but an infinite number of forms of f{t) consistent with this value would be consistent with any arbitrary value of 5 at a further moment That is, even if all the observed values agree with

(1) exactly, deductive logic can say nothing whatever about the value of s at any other time. An infinite number of laws agree with previous experience, and an infinite number that have agreed with previous ex> perience will inevitably be wrong in the next instance. What the applied mathematician does, in fact, is to select one form out of this infinity; and his reason for doing so has nothing whatever to do with traditional logic. He chooses the simplest. This is actually an understatement of the case; because in general the observations will not agree with (1)

4

FUNDAMENTAL NOTIONS

Chap. I

exactly, a polynomial of n terms can still be found that will agree exactly with the observed values at times /p..., and yet the form (1) may still be asserted. Similar considerations apply to any quantitative law. The further discussion of this matter must be reserved till we come to significance tests. We need notice at the moment only that the choice of the simplest law that fits the facts is an essential part of procedure in applied mathematics, and cannot be justified by the methods of deductive logic. It is, however, rarely stated, and when it is stated it is usually in a manner suggesting that it is something to be ashamed of. We may recall the words of Brutus.

But ’tis a common proof That lowliness is young ambition’s ladder,

Wlieroto the climber upwards turns his face ;

But when he once attains the upmost round,

He then imto the ladder turns his back,

Looks in the clouds, scorning thf' bas(3 degrees By which he did ascend.

It is asserted, for instance, that the choice of the simplest law is purely a matter of economy of description or thought, and has nothing to do with any reason for believing the law. No reason in deductive logic, certainly; but the question is, Does deductive logic contain the whole of reason? It does give economy of description of past experience, but is it unreasonable to be interested in future experience ? Do we make predictions merely because those predictions are the easiest to make ? Does the Nautical Almanac Office laboriously work out the positions of the planets by means of a complicated set of tables based on the law of gravitation and previous observations, merely for convenience, when it might much more easily guess them? Do sailors trust the safety of their ships to the accuracy of these predictions for the same reason? Does a town install a new tramway system, with expensive plant and much preliminary consultation with engineers, with no more reason to suppose that the trams will move than that the laws of electromagnetic induction are a saving of trouble? I do not believe for a moment that anybody will answer any of these questions in the affirmative; but an affirmative answer is implied by the assertion that is still frequently made, that the choice of the simplest law is merely a matter of convention. I say, on the contrary, that the simplest law is chosen because it is the most likely to give correct predictions; that the choice is based on a reasonable degree of belief; and that the fact that deductive logic provides no explanation of the choice of the simplest law is an absolute proof that deductive logic is grossly inadequate to

§1.0 FUNDAMENTAL NOTIONS 6

cover scientific and practical requirements. It is sometimes said, again, that the trust in the simple law is a peculiarity of human psychology; a different type of being might behave differently. Well, I see no point whatever in discussing at length whether the human mind is any use; it is not a perfect reasoning instrument, but it is the only one we have. Deductive logic itself could never be known without the human mind. If anybody rejects the human mind and then holds that he is construct- ing valid arguments, he is contradicting himself; if he holds that human minds other than his own are useless, and then hopes to convince them by argument, he is again contradicting himself. A critic is himself using inductive inference when he expects his words to convey tiie same meaning to his audience as they do to himself, since the meanings of words are learned first by noting the correspondence between things and the sounds uttered by other l)eople, and then applied in new instances. On the face of it, it would appear that a general state- ment that something accepted by the bulk of mankind is intrinsically nonsense requires much more to support it than a mere declaration.

Many attempts have been made, while accepting induction, to claim that it can be reduced in some way to deduction. Bertrand Russell has remarked that induction is either disguised deduction or a mere method of making plausible guesses,! the former sense we must look for some general principle, which states a set of possible alternatives; then observations are used to show that all but one of these are wrong, and the survivor is held to be deductively demonstrated. Such an attitude has been widely advocated. On it I quote Professor C. D. Broad, t

‘The usual view of the logic books seems to bo that inductive arguments are really syllogisms with propositions summing up the relevant observations as minors, and a common major consisting of some universal proposition about nature. If this were true it ought to be easy enough to tind the missing major, and the singular obscurity in which it is enshrouded would be cjuite inexplicable. It is reverently referred to by inductive logicians as the Uniformity of Nature; but, as it is either never stated at all or stated in such terms that it could not possibly do what is required of it, it appears to be the inductive equivalent of Mrs. Gamp’s mysterious friend, and might be more appropriately named Major Harris.

t Principles of Mathematics, p. 360, He said, at the Aristotelian Society summer meeting in 1938, that this remark has been too much quoted. I therefore offer apologies for quoting it again. He has also remarked that the inductive philosophers of Central Africa formerly held the view that all men were black. My comment would be that the deductive ones, if there were any, did not hold that there wore any men, black, white, or yellow.

t Mind, 29, 1920, 11.

6

FUNDAMENTAL NOTIONS

Chap. I

‘It is in fact easy to prove that this whole way of looking at inductive argu- ments is mistaken. On this view they are all syllogisms with a common major. Now their minors are propositions summing up the relevant observations. If the observations have been carefully made the minors are practically certain. Hence, if this theory were true, the conclusions of all inductive arguments in which the observations were equally carefully made would bo equally probable. For what could vary the probabilities ? Not the major, which is common to all of them. Not the minors, which by hypothesis are equally certain. Not the mode of reasoning, which is syllogistic in each case. But the result is preposterous, and is enough to refute the theory which loads to it.’

Attempts have been made recently to supply the missing major by several modern physicists, notably Sir Arthur Eddington and Professor E. A. Milne. But their general principles and their results differ even within the very limited field of knowledge where they have been applied. How is a person with less penetration to know which is right, if any? Only by comj)aring the results with observation: and then his reason for believing the survivor to be likely to give t)u^ right results in future is inductive. I am not denying that one of them may liave got the right results. But I reject the statement that any of them can be said to be certainly right as a matter of pure logic, independently of experience; and I gravely doubt whether any of them could have been thought of at all had the authors been unaware of the vast amount of previous work that had led to the establishment by inductive methods of the laws that they set out to explain These attempts, though they appear to avoid Broad’s objection, do so only within a limited range, and it is doubtful whether such an attempt is w orth making if it can at best achieve a partial success, when induction can cover the whole field without supposing that special rules hold in certain subjects.

I should maintain (with N. R. Campbell, who saysj that a physicist would be more likely to interchange the two terms in Russell’s state- ment) that a great deal of what passes for deduction is really disguised induction, and that even some of the postulates of Principia Mathe- matica are adopted on inductive grounds (which, incidentally, are false).

Two attempts at a justification of induction, still sometimes made, are as follows. (1) Induction has worked in the past; therefore it will work in the future. It is obvious that this is itself an inductive inference and involves the same problems in a more complicated way. (2) The struggle for existence would favour members with the ability to predict correctly the consequences of their actions. Consequently the fact that man has survived implies that he has this ability (and presumably

f Physics f The Ele7nent8^ 1920, 9.

FUNDAMENTAL NOTIONS

7

§ 1.0

Amoeba has too). But the belief that there is a struggle for existence and that it favours particular types is based on induction. Both argu- ments replace the original question by another as difficult or more so, and take no effective step towards a solution.

Karl Pearson*}* writes as follows:

‘Now this is the peculiarity of scientific method, that when once it has become a habit of mind, that mind converts all facts whatsoever into science. The field of science is unlimited ; its material is endless, every group of natural phenomena, every phase of social life, every stage of past or present development is material for science. The unity of all science consists alone in its method, not in its material. The man who classifies facts of any kind whatever, who sees their mutual relation and describes their sequences, is applying the scientific method and is a man of science. The facts may belong to the past history of mankind, to the social statistics of our great cities, to the atmosphere of the most distant stars, to the digestive organs of a worm, or to the life of a scarcely visible bacillus. It is not the facts themselves which form science, but the methods by which they are dealt with.’

Here, in a few sentences, Pearson sets our problem. The italics are his. He makes a clear distinction between method and material. No matter what the subject-matter, the fundamental principles of the method must be the same. There must be a uniform standard of validity for all hypotheses, irrespective of the subject. Different laws may hold in different subjects, but they must be tested by the same criteria other- wise we have no guarantee that our decisions w ill be those warranted by the data and not merely the result of inadequate analysis or of believing what we w^ant to believe. An adequate theory of induction must satisfy two conditions. First, it must provide a general method; secondly, the principles of the method must not of themselves say any- thing about the world. If the rules are not general, we shall have different standards of validity in different subjects, or different standards for one’s own hypotheses and somebody else’s. If the rules of themselves say anything about the world, they wdll make empirical statements independently of observational evidence, and thereby limit the scope of what we can find out by observation. If there are such limits, they must be inferred from observation; we must not assert them in advance.

We must notice at the outset that induction is more general than deduction. The answers given by the latter are limited to a simple ‘yes’, ‘no’, or ‘it doesn’t follow’. Inductive logic must split up the last alternative, which is of no interest to deductive logic, into a number of others, and say wffiich of them it is most reasonable to believe on

t The Grammar of Science, 1802. P. 16 of Everyman edition, 1938.

8

FUNDAMENTAL NOTIONS

Chap. I

the evidence available. Complete proof and disproof are merely the extreme cases. Any inductive inference involves in its very nature the possibility that the alternative chosen as the most likely may in fact be wrong. Exceptions are always possible, and if a theory does not provide for them it will be claiming to be deductive when it cannot be. On account of this extra generality, induction mxist involve postulates not included in deduction. Our problem is to state these postulates. It is important to notice that they cannot be proved by deductive logic. If they could, induction would be reduced to deduction, which is impossible. Equally they are not empirical generalizations; for in- duction would be needed to make them and the argument would be circular. We must in fact distinguish the general rules of the theory from the empirical content. The general rules are a priori propositions, accepted independently of experience, and making by themselves no statement about experience. Induction is the application of the rules to observational data.

Our object, in short, is not to prove induction; it is to tidy it up. Even among professional statisticians there are considerable differences about the best way of treating the same problem, and, I think, all statisticians would reject some methods habitual in some branches of physics. The question is whether we can construct a general method, the acceptance of which would avoid these differences or at least reduce them.

1.1. The test of the general rules, then, is not any sort of proof. This is no objection because the primitive propositions of deductive logic cannot be proved either. All that can be done is to state a set of hypotheses, as plausible as possible, and see where they lead us. The fullest development of deductive logic and of the foundations of mathe- matics is that of Principia Mathematica, which starts with a number of primitive propositions taken as axioms; if the conclusions are accepted, that is because we are willing to accept the axioms, not because the latter are proved. The same applies, or used to apply, to Euclid. We must not hope to prove our primitive propositions when this is the position in pure mathematics itself. But we have rules to guide us in stating them, largely suggested by the procedure of logicians and pure mathematicians .

1. All hypotheses used must be explicitly stated, and the conclusions must follow from the hypotheses.

2. The theory must be self-consistent; that is, it must not be possible

§ 1.1 FUNDAMENTAL NOTIONS 9

to derive contradictory conclusions from the postulates and any given set of observational data.

3. Any rule given must be applicable in practice. A definition is useless unless the thing defined can be recognized in terms of the definition when it occurs. The existence of a thing or the estimate of a quantity must not involve an impossible experiment.

4. The theory must provide explicitly for the possibility that infer- ences made by it may turn out to be wrong. A law may contain adjustable parameters, which may be wrongly estimated, or the law itself may be afterwards found to need modification. It is a fact that revision of scientific laws has often been found necessary in order to take account of new information the relativity and quantum theories providing conspicuous instances and there is no conclusive reason to suppose that any of our present laws are final. But we do accept inductive inference in some sense; we have a certain amount of con- fidence that it will be right in any particular case, though this confidence does not amount to logical certainty.

5. The theory must not deny any empirical proposition a priori \ any precisely stated empirical proposition must be formally capable of being accepted, in the sense of the last rule, given a moderate amount of relevant evidence.

These five rules are essential. The first two impose on inductive logic criteria already required in pure mathematics. The third and fifth enforce the distinction between a priori and empirical propositions; if an existence depends on an inapplicable definition we must either find an applicable one, treat the existence as an empirical proposition requiring test, or abandon it. The fourth states the distinction between induction and deduction. The fifth makes Pearson’s distinction be- tween material and method explicit, and involves the definite rejection of attempts to derive empirically verifiable propositions from general principles adopted independently of experience.

The following rules also serve as useful guides.

6. The number of postulates should be reduced to a minimum. This is done for deductive logic in Principia, though many theorems proved there appear to be as obvious intuitively as the postulates. The motive for not accepting other obvious propositions as postulates is partly artistic. But we cannot regard the human mind as a perfect reasoner, and a reduction of the number of postulates affords a check on the consistency of different propositions, any of which we might be ready to accept by itself. This is still more needed in induction, since the

10

FUNDAMENTAL NOTIONS

Chap. I

beliefs often accepted as intuitiv^ely certain are more numerous, and, I believT, some of them are definitely inconsistent, while others are not primitive propositions but inductive inferences. If they are, they can- not, of course, be asserted as certain, but they may be asserted with so high a probability that there will be little difference in practice.

7. While we do not regard the human mind as a perfect reasoner, we must accept it as a useful one and the only one available. The theory need not represent actual thought-processes in detail, but should agree with them in outline. We are not limited to considering only the thought-processes that people describe to us. It often happens that their behavdour is a better criterion of their inductive processes than their arguments. If a result is alleged to be obtained by arguments that are certainly wrong, it does not follow that the result is Avrong, since it may have been obtained by a rough inductive process that the author thinks it undesirable or unnecessary to state on account of the traditional insistence on deduction as the only valid reasoning. I disagree utterly with many arguments produced by the chief current schools of statistics, but I rarely differ .seriously from the conclusions; their practice is far better than their precept. 1 should say that this is the result of common sense emerging in spite of the deficiencies of mathematical teaching. The theory must provide criteria for testing the chief types of scientific laAv that have actually been suggested or asserted. Any such law must be taken seriously in the sense that it can be asserted with confidence on a moderate amount of evidence. The fact that simple laws are often asserted Avill, on this criterion, require us to say that in any particular instance some simple law is quite likely to be true.

8. In view of the greater complexity of induction, we cannot hope to develop it more thoroughly than deduction. We shall therefore take it as a rule that an objection carries no weight if an analogous objection would invalidate part of generally accepted pure mathematics. I do not wish to insist on any particular justification of pure mathematics, since authorities on its foundations are far from being agreed among themselves. In Prineijna much of higher mathematics, including the whole theory of the continuous variable, rests on the axioms of infinity and reducibility, which are rejected by Hilbert. F. P. Ramsey rejects the axiom of reducibility, while declaring that the multiplicative axiom, properly stated, is the most evident tautology, though Whitehead and Russell express much doubt about it and carefully separate propositions that depend on it from those that can be proved without it. I should

§1.1

FUNDAMENTAL NOTIONS

11

go further and say that the proof of the existence of numbers, according to the Princi'pia definition of number, depends on the postulate that all individuals are permanent, which is an empirical proposition, and a false one, and should not be made part of a deductive logic. But we do not need such a proof for our purposes. It is enough that pure mathematics should be consistent. If the postulate could hold in 807ne world, even if it was not the actual world, that would be enough to establish consistency. Then the derivation of ordinary mathematics from the postulates of Principia can be regarded as a proof of its con- sistency. But the justification of all the justifications seems to be that they lead to ordinary pure mathematics in the end; I shall assume that the latter has validity irrespective of any particular justification.

The above principles will strike many readers as platitudes; and if they do I shall not object. But they require the rejection of several principles accepted as fundamental in other theories. They rule out, in the first place, any definition of probability that attempts to define probability in terms of infinite sets of possible observations, for we cannot in practice make an infinite number of observations. The Venn limit, the hypothetical infinite population of Fisher, and the ensemble of Willard Gibbs are useless to us by rule 3. Though many accepted results appear to be based on these definitions, a closer analysis shows that further hypotheses are required before any results are obtained, and these hypotheses are not stated. In fact, no 'objective’ definition of probability in terms of actual or possible observations, or possible properties of the world, is admissible. For, if we made anything in our fundamental principles depend on observations or on the structure of the world, we should have to say either (1) that the observations we can make, and the structure of the world, are initially unknown; then we cannot know our fundamental principles, and we have no possible starting-point; or (2) that we know a priori something about observa- tions or the structure of the world, and this is illegitimate by rule 5. Attempts to use the latter principle will superpose our preconceived notions of what is objective on the entire system, whereas, if objectivity has any meaning at all, our aim must be to find out what is objective by means of observations. To try to give objective definitions at the start will at best produce a circular argument, may lead to contradic- tions, and in any case will make the whole scheme subjective beyond hope of recovery. We must not rule out any empirical proposition a priori] we must provide a system that will enable us to test it when occasion arises, and this requires a completely comprehensive formal scheme.

12

FUNDAMENTAL NOTIONS

Chap. I

We must also reject what is variously called the principle of causality, determinism, or the uniformity of nature, in any such form as ‘Precisely similar antecedents lead to precisely similar consequences’. No two sets of antecedents are ever identical; they must differ at least in time and position. But even if we decide to regard time and position as irrelevant (which may be true, but has no justification in pure logic) the antecedents are never identical. In fact, determinists usually recog- nize this verbally and try to save the principle by restating it in some such form as: ‘In precisely tlie same circumstances very similar things can be observed, or very similar things can usually be observed. ’j If ‘precisely the same’ is intended to be a matter of absolute trutli, we cannot achieve it. Astronomy is usually considered a science, but the planets have never even approximately repeated tlieir positions since astronomy began. The principle gives us no means of inferring the accelerations at a single instant, and is utterly useless. Further, if it was to be any use we should have to know at any application that the entire condition of the world w as the same as in some previous instance. This is never satisfied in tiie most carefully controlled experimental conditions. The most that can be done is to make those conditions the same that we believe to be relevant ‘the same’ can never in practice mean more than ‘the same as far as we know^’, and usually means a great deal less. The question then arises, How do we know tliat the neglected variables are irrelevant? Only by actually allowing them to vary and verifying that there is no associated variation in the result; but this requires the use of significance tests, a theory of which must therefore be given before there is any application of the principle, and when it is given it is found that the principle is no longer needed and can be omitted by rule fi. It may conceivably be true in some sense, though nobody has succeeded in stating clearly what this sense is. But what is quite certain is that it is useless.

Causality, as used in applied mathematics, has a more general form, such as: ‘Physical laws are expressible by mathematical equations, possibly connecting continuous variables, such that in any case, given a finite number of parameters, some variable or set of variables that appears in the equations is uniquely determined in terms of the others.’ This does not require that the values of the relevant parameters should be actually repeated; it is possible for an electrical engineer to predict the performance of a dynamo without there having already been some exactly similar dynamo. The equations, which we call law s, are inferred t W. H. George, The Scientist in Action^ 1936, p. 48.

§1.1

FUNDAMENTAL NOTIONS

13

from previous instances and then applied to instances where the relevant quantities are different. This form permits astronomical prediction. But it still leaves the questions 'How do we know that no other para- meters than those stated are needed?’, ‘How do we know that we need consider no variables as relevant other than those mentioned explicitly in the laws?’, and ‘Why do we believe the laws themselves?’ It is only after these questions have been answered that we can make any actual application of the principle, and the ]>rinciple is useless until we have attended to the epistemological problems. Further, the principle happens to be false for quantitative observations. It is not true that observed results agree exactly with the predictions made by the laws actually used. The most that the laws do is to predict a variation that accounts for the greater part of the observed variation ; it never accounts for the whole. The balance is called ‘error’ and usually quickly for- gotten or altogether disregarded in physical writings, but its existence compels us to say that the laws of applied mathematics do not express the whole of the variation. Their justification cannot be exact mathe- matical agreement, but only a partial one depending on what fraction of the observed variation in one quantity is accounted for by the variations of the others. The phenomenon of error is often dealt with by a suggestion of various minor variations that might alter the measurements, but this is no answer. An exact quantitative prediction could never be made, even if such a suggestion was true, unless we knew in each individual case the actual amounts of the minor varia- tions, and we never do. If we did w'e should allow^ for them and obtain a still closer agreement; but the fact remains that in practice, however fully w^e take small variations into account, we never get exact agree- ment. A physical law, for practical use, cannot be merely a statement of exact predictions; if it was it would invariably be wrong and w ould be rejected at the next trial. Quantitative prediction must always be prediction within a margin of uncertainty; the amount of this margin will he different in different cases, but for a law' to be of any use it must state the margin explicitly. The outstanding variation, for prac- tical application, is as essential a part of the law^ as the predicted variation is, and a valid statement of the law^ must express it. But in any individual case this outstanding variation is not known. We know only something about its possible range of values, not what the actual value will be. Hence a physical law is not an exact prediction, but a state- ment of the relative probabilities of variations of different amounts. It is only in this form that we can avoid rejecting causality altogether as false,

14

FUNDAMENTAL NOTIONS

Chap. I

or as inapplicable under rule 3 ; but a statement of ignorance of the individual errors has become an essential part of it, and we must recognize that the physical law itself, if it is to be of any use, must have an epistemological content.

The impossibility of exact prediction Jias recentl}' been forced on the attention of physicists by Heisenberg's Uncertainty Principle. It is remarkable, considering that the phenomenon of errors of observation was discussed by Laplace and Gauss, that there should still have been any physicists that thought that actual observations were exactly pre- dictable; yet attempts to evade the principle have shown that many exist. The principle is actually no new uncertainty. What Heisenberg has done is to consider the most refined types of observation that modern physics suggests might be possible, and to obtain a lower limit to the uncertainty; but it is much smaller than the old uncertainty, which was never neglected except by misplaced optimism. The exist- ence of errors of observation seems to have escaped the attention of many philosophers that have discussed the uncertainty principle; this is perhaps because they tend to get their notions of physics from popular writings, and not from works on the combination of observations. Their criticisms of popular physics, mostly valid as far as they go, would gain enormously in force if they attended to what we knew about errors before Heisenberg. f

The word error is liable to be interpreted in some ethical sense, but its scientific meaning is closely connected with the original one. Latin errare, in its original sense, means to wander, not to sin or to make a mistake. The meaning occurs in ‘knight -errant'. The error means simply the outstanding variation after we have done our best to inter- pret the whole variation.

The criterion of universal assent, stated by Dr. N. R. Campbell and by Professor H. Dingle in his Science and Human Experience (but abandoned in his Through Science to Philosophy), must also be rejected

t Professor L. S. Stebbing {Philosophy and the Physicists, 1938, p. 198) remarks: ‘There can be no doubt at all that precise predictions concerning the behaviour of macroscopic bodies are made and are exactly verified within the limits of experimental error.’ Without the saving phase at the end the statement is intelligible, and false. With it, it is meaningless. The severe criticism of much in modern physics contained in this book is, in my opinion, thoroughly justified, but the later parts lose much of their point through inattention to the problem of errors of observation. Some philo- sophers, however, have seen the point quite clearly. For instsmee. Professor J. H. Muirhead {The Elements of Ethics, 1910, pp. 37-8) states: ‘The truth is that what is called a natural law is itself not so much a statement of fact as of a standard or type to which facts have been foimd more or less to approximate. This is true ev'pn in inorganic nature.’ I am indebted to Mr. John Bradley for the reference.

§1.1 FUNDAMENTAL NOTIONS 16

by rule 3. This criterion requires general acceptance of a principle before it can be adopted. But it is impossible to ask everybody’s con- sent before one believes anything; and if ‘everybody’ is replaced by ‘everybody qualified to judge’, we cannot apply the criterion until we know who is qualified, and even then it is liable to happen that only a small fraction of the people capable of expressing an opinion on a scientific paper read it at all, and few even of those do express any. Campbell lays much stress on a physicist’s characteristic intuition, j* which apparently enables him always to guess right. But if there is any such intuition there is no need for the criterion of general agree- ment or for any other. The need for some general criterion is that even among those apparently qualified to judge there are often serious differences of opinion about the proper interpretation of the same facts; what we need is an impersonal criterion that will enable an individual to see whether, in any particular instance, he is following the rules that other people follow and that he himself follows in other instances.

1.2. The chief constructive rule is 4. It declares that there is a valid primitive idea expressing the degree of confidence that we may reason- ably have in a proposition, even though we may not be able to give either a deductive proof or a disproof of it. In extreme cases it may be a mere statement of ignorance. We need to express its rules. One obvious one (though it is very commonly overlooked) is that it depends both on the proposition considered and on the data in relation to which it is considered. Suppose that I know that Smith is an Englishman, but otherwise know nothing particular about him. He is very likely, on that evidence, to have a blue right eye. But suppose that I am informed that his left eye is brown the probability is changed com- pletely. This is a trivial case, but the principle in it constitutes most of our subject-matter. It is a fact that our degrees of confidence in a proposition habitually change when we make new observations or new evidence is communicated to us by somebody else, and this change constitutes the essential feature of all learning from experience. We must therefore be able to express it. Our fundamental idea will not be simply the probability of a proposition j), but the probability of p on data g. Omission to recognize that a probability is a function of two arguments, both propositions, is responsible for a large number of serious mistakes; in some hands it has led to correct results, but at the t Arisiot. Soc. Suppl. vol. 17, 1938, 122.

16

FUNDAMENTAL NOTIONS

Chap. I

cost of omitting to state essential liypotheses and giving a delusive appearance of simplicity to what are really very difficult arguments. It is no more valid to speak of the probability of a proposition without stating the data than it would be to speak of the value of x-\-y for given x, irrespective of the value of y.

We can now proceed on rule 7. It is generally believed that proba- bilities are orderable: that is, that if p, q, r are three propositions, the statement ‘on data p, q is more probable than r’ has a meaning. In actual cases people may disagree about which is the more probable, and it is sometimes said that this implies that the statement has no meaning. But the differences may have other explanations; (1) Tlie commonest is that the probabilities are on different data, one person having relevant information not available to the other, and we liave made it an essential point that the probability de})ends on the data. The conclusion to draw in such a case is that, if people argue without telling each other what relevant information they have, they are wasting their time. (2) The estimates may be wrong. It is perfectly possible to get a wrong answer in pure mathematics, so that by rule 8 this is no objection. In this case, where the probability is often a mere guess, we cannot expect the answer to be right, tliough it inay be and often is a fair approximation. (3) The wish may be father to the thought. But perhaps this also has an analogue in pure mathematics, if w e con- sider the number of fallacious methods of squaring the circle and proving Fermat’s last theorem that have been given, merely because people wanted n to be an algebraic or rational number or the theorem to be true. In any case alternative hypotheses are open to the same objection, on the one hand, that they depend on a wish to have a wholly deductive system and to avoid the explicit statement of the fact that scientific inferences are not certain; or, on the other, that the statement that there is a most probable alternative on given data may curtail their freedom to believe another wffien they find it more pleasant. I think that these reasons account for all the apparent differences, but they are not fundamental. Even if people disagree about which is the more probable alternative, they agree that the comparison has a meaning. We shall assume that this is right. The meaning, however, is not a statement about the external world; it is a relation of inductive logic. Our primitive notion, then, is that cf the relation ‘given p, q is more probable than r\ where p, q, and r are three propositions. If this is satisfied in a particular instance, we say that r is less probable than q, given p; this is the definition of less probable. If given <1 i^^ neither

FUNDAMENTAL NOTIONS

17

§ 1.2

more nor less probable than r, q and r are equally probable, given p. Then our first axiom is

Axiom 1. Given p, q is either more, equally, or less probable than r, and no two of these alternatives can be true.

This axiom may be called that of the comparability of probabilities. In Scientific Inference I took it in a more general form, assuming that the probabilities of propositions on different data can be compared. But this appears to be unnecessary, because it is found that the com- parability of probabilities on different data, whenever it arises in practice, is proved in the course of the work and needs no special axiom. The fundamental relation is transitive; we express this as follows.

Axiom 2. If p^ q, r, s are four propositions, and, given p, q is more probable than r ayid r is more q^vobable than s, then, given p, q is more probable than s.

The extreme degrees of probability are certainty and impossibility. These lead to

Axiom 3. All proj^ositions deducible from a proposition p have the same probability on data p; ami all projyositions inconsistent with p have the same probability on data p.

We need this axiom to ensure consistency with deductive logic in cases that can be treated by both methods. We are trying to construct an extended logic, of which deductive logic will be a part, not to intro- duce an ambiguity in cases where deductive logic already gives definite answers. I shall often speak of ‘certainty on data p’ and ‘impossibility on data p\ These do not refer to the mental certainty of any particular individual, but to the relations of deductive logic expressed by is deducible from p' and ‘not-g is deducible from p\ In G. E. Mooi'e’s terminology, we may read the former as ‘p entails q\ In consequence of our rule 5, we shall never have ‘p entails g' ifp is merely the general rules of the theory and q is an empirical proposition.

Actually I shall take 'entails’ in a slightly extended sense; in some usages it would be held that p is not deducible from p, or from p and q together. Some shortening of the writing is achieved if we agree to define ‘p entails q^ as meaning either ‘g is deducible from p’ or 'q is identical with p’ or ' (7 is identical with some proposition asserted in p’. This avoids the need for special attention to trivial cases.

We also need the following axiom.

Axiom 4. If, given p, q and q' cannot both be true, and if, given p,

3595.58 n

18

FUNDAMENTAL NOTIONS

Chap. I

r and r' cannot both be true, and if, given p, q and r are equally probable and q' and r' are equally probable, then, given p,'q or q'" and r or r'’ are equally probable.

At this stage it is desirable for clearness to introduce the following notations and terminologies, mainly from Frincipia Mathematica. p means ‘not-^’; that is, p is false.

p,q means 'p and q'; that is, p and q are both true.

p ^ q means 'p ov q'\ that is, at least one of p and q is true.

These notations may be combined, dots being used as brackets. Thus :p.q means 'p and q is not true’; that is, at least one of jo and q is false, which is equivalent to ^ p,y. ^ q. But

^ p.q means 'p is false and q is true’, which is not the same pro- position. The rule is that a set of dots represents a bracket, the com- pletion of the bracket being either the next equal set of dots or the end of the expression. Dots may be omitted in joint assertions where no ambiguity can arise.

The^’om^ assertion or conjunction ot p and q is the proposition p .q, and the joint assertion of p, q, r, s,... L the proposition p.q,r ,s...\ that is, that p, q, r, s,... are all true. The joint assertion is also called the logical product.

The disjunction of p and q is the proposition pvq; the disjunction of q, r, s is the proposition p v q v r ^ s, that is, at least one of p, q, r, s is true. The disjunction is also called the logical sum.

A set of propositions q^ {i 1 to /i) are said to be exclusive on data p if not more than one of them can be true on data p, that is, if p entails all the disjunctions qi^ qu when i ^ k.

A set of propositions q, r, s are said to be exhaustive on data p if at least one of them must be true on data p; that is, if p entails the dis- junction qy r V s.

It is possible for a set of alternatives to be both exclusive and exhaustive. For instance, a finite class must have some number n; then the propositions n = 0, 1, 2, 3,... must include one true proposi- tion, but cannot contain more than one.

Then Axiom 4 will read:

If q and q' are exclusive, and r and r' are exclusive, on data p, and if, given p, q and r are equally probable and qf and F are equally probable, then, given p, qy q' and ryr' are equally probable.

An immediate extension, obtained by successive applications of this axiom, is:

§1-2

FUNDAMENTAL NOTIONS

19

Theorem 1. IJ q^ are exclusive, and r^, are exclusive,

on dfOta p, and if, given p, the proqyositions q^ and r^, q^ and r^,,..,q^ and r^ are eqtially probable in pairs, then given p, q^^ q^-- qn ^ ^2 ^

are equally probable.

It will be noticed that we have not yet assumed that probabilities can be expressed by numbers. I do not think that the introduction of numbers is strictly necessary to the further development; but it has the enormous advantage that it permits us to use mathematical technique. Without it, while we might obtain a set of propositions that would have the same meanings, their expression would be much more cumbrous. The actual introduction of numbers is done by conventions, the nature of which is essentially linguistic.

f'oNVENTiON 1. Wc assign the larger number on given data to the more probable proposition {and therefore equal numbers to equally probable propositions).

Convention 2. If, given p, q a7id q' are exclusive, then the number assigned on data p to' q or q" is the sum of those assigned to q and to q\

It is important to notice the meaning of a convention. It is neither an axiom nor a theorem. It is merely a rule introduced for convenience, and it has the property that other rules would give the same results. W. E. Johnson remarks that a convention is properly expressed in the imperative mood. An instance is the use of rectangular or polar coordi- nates in Euclidean geometry. The distance between two points is the fundamental idea, and all propositions can be stated as relations be- tween distances. Any proposition in rectangular coordinates can be translated into polar coordinates, or vice versa, and both expressions would give the same results if translated into propositions about distances. It is purely a matter of convenience which we choose in a particular case. The choice of a unit is always a convention. But care is needed in introducing conventions; some postulate of consistency about the fundamental ideas is liable to be hidden. It is quite easy to define an equilateral right-angled plane triangle, but that does not make such a triangle possible. In this case Convention 1 specifies what order the numbers are to be arranged in. Numbers can be arranged in an order, and so can probabilities, by Axioms 1 and 2. The relation ‘greater than’ between numbers is transitive, and so is the relation ‘more probable than’ between propositions on the same data. There- fore it is possible to assign numbers by Convention 1, so that the order of increasing degrees of belief will be the order of increasing number.

20 FUNDAMENTAL NOTIONS Chap. I

So far we need no new axiom; but we shall need the axiom that there are enough numbers for our purpose.

Axiom 5. The set of possible probabilities on given datu, ordered in terms of the relation ^more probable than\ can be put into one-one corre- spondence with a, set of real numbers in increasing order.

The need for such an axiom was pointed out by an American reviewer of Scientific Inference, He remarked that if we take a series of number pairs ^ (a,,, 6„) and make it a rule that Uj. is to be placed after if but that if a^. is to be placed after if b^ > then

the axiom that the can be placed in an order will hold, but if and b^^ can each take a continuous series of values it v ill be impossible to establish a one-one correspondence between the pairs and a single continuous series without deranging the order.

Convention 2 and Axiom 4 wdll imply that, if we have two pairs of exclusive propositions with the same probabilities on the same data, the numbers chosen to correspond to their disjunctions wdll be the same. The extension to disjunctions of several propositions is justi- fied by Theorem 1. We shall alw^ays, on given data, associate the same numbers with propositions entailed or contradicted by the data; this is justified by Axiom 3. The assessment of numbers in the way suggested is therefore consistent with our axioms. We can now intro- duce the formal notation , v

P{q I p)

for the number associated with the probability of the proposition q on data p; it may be read 'the probability of q given p' provided that we remember that the number is not in fact the probability, but merely a representation of it in terms of a pair of conventions. The probability, strictly, is the reasonable degree of confidence and is not identical with the number used to express it. The relation is that between Mr. Smith and his name ‘Mr. Smith'. A sentence containing the words 'Mr. Smith’ may correspond to, and identify, a fact about Mr. Smith. But Mr. Smith himself does not occur in the sentence.! In this notation, the properties of numbers will now replace Axiom 1 ; Axiom 2 is restated ‘if P{q Ip) > P(r I p), and P(r\p) > P(s |p), then P{q \p) > P{s \p)\ which is a mere mathematical implication, since all the expressions are numbers. Axiom 3 will require us to decide what numbers to associate with certainty and impossibility. We have

Theorem 2. If p is consistent with the general rules, and p entails ~ g, then P(q \p) ^ 0.

t Cf. R. Carnap, The Logical Syntax of Language.

§1.2 FUNDAMENTAL NOTIONS 21

For let q and r be any two propositions, both impossible on data p. Then (Ax. 3) if a is the number associated with impossibility on data p,

P{q ip) =- P(r Ip) P{q v r |p) a

since q, r, and g v r are all impossible propositions on data p and must be associated with the same number. But qr is impossible on data p; hence, by definition, q and r are exclusive on datap, and (Conv. 2)

P{qwr\p) -- P{q\p) + P{r\p) - 2a;

whence a ~ 0. Therefore all probability numbers are > 0, by Con- vention 1.

As we have not assumed the comparability of probabilities on dif- ferent data, attention is needed to the possible forms tliat can be substituted for q and r, given p. If p is a purely a priori proposition, it can never entail an empirical one. Hence, if p stands for our general rules, the admissible values for q and r must be false a priori proposi- tions, such as 2 = 1 and 3^2. Since such propositions can be stated the theorem follows. If p is empirical, then p is an admissible value for both q and r. Or, since we are maintaining the same general prin- ciples throughout, we may remember that in practice if p is empirical and we denote our general principles by h, then any set of data that actually occurs and includes an empirical proposition will be of the form ph. Then for q and r we may still substitute false a priori pro- positions, which will be impossible on data ph. Hence it is always possible to assign q and r so as to satisfy the conditions stated in the proof.

Convention 3. If p entails q, then P{q |p) 1.

This is the rule generally adopted; but there are cases where we wish to express ignorance over an infinite range of values of a quantity, and it is then convenient to express certainty that the quantity lies in that range by oo, in order to keep ratios for finite ranges determinate. None of our axioms so far has stated that we must always express certainty by the same number on different data, merely that we must on the same data; but with this exception it is convenient to do so.

The converse of Theorem 2 would be: ‘If P{q jp) ~ 0, then p entails ~ g.’ This is false if we use Convention 3. For instance, a continuous variable may be equally likely to have any value between 0 and 1. Then the probabihty that it is exactly J is 0, but | is not an impossible value. There would be no point in making certainty correspond to infinity in such a case, for it would make the probability infinite for

22 FUNDAMENTAL NOTIONS Chap. I

any finite range. It turns out that we have no occasion to use the converse of Theorem 2.

Axiom 6, If pq eyitails r, then P(qr \ p) ~ F(<q \p).

In other words, given p throughout, we may consider whether q is false or true. If q is false, then qr is false. If q is true, then, since pq entails r, r is also true and therefore qr is true. Similarly, if qr is true it entails q, and if qr is false q must be false on data p, since if it was true qr would be true. Thus it is impossible, given />, that either q or qr should be true without the other. This is an extension of Axiom 3 and is necessary to enable us to take over a further set of rules sug- gested by deductive logic, and to say that all equivalent propositions have the same probability on given data.

Theorem 3. If q and r are equivalent in the sense that each entails the other, then each entails qr, and the probabilities of q and r on any data must be equal. Shnilarly, if pq entails r, and pr entails q, P{q |p) - P{r\p), since both are equal to P{qr | p).

An immediate corollary is

Theorem 4. P(q |p) = P{qr \p)-\-P{q. r-^ r \p).

For qr and q. ^ r are exclusive, and the sum of their probabilities on any data is the probability of qr:v:q. ^r (Conv. 2). But q entails

this proposition, and also, if either q and r are both true or q is true

and r false, q is true in any case. Hence tho propositions q and qr:w:q, ^ r are equivalent, and the theorem follows by Theorem 3.

It follows further that P(q \ p) ^ P{^qr j p), since P^q. ^ \ p) cannot be negative. Also, if we write q v r for q, we have

P{q V r Ip) P(q v r:r |p)+P(^ y r: r |p) (Th. 4)

and ^ V r : r is equivalent to r, and q y r: r to q. r. Hence

P(qyr\p) ^ P(r\p).

Theorem 5. If q and r are two propositions, not necessarily exclusive on data p,

P{q\p)+P(.'r\p) = P(qy r\p) + P{qr\p).

For the propositions qr, q. ^ r, q.r, q. ^ r are exclusive; and q is equivalent to the disjunction of qr and q. r, and r to the dis- junction of gr and ^ q.r. Hence the left side of the equation is equal to

2P{qr\p)+P(q. r \p)-\-P{-^ q.r \p) (Th. 4).

.Also g V r is equivalent to the disjunction of qr, q. r, and q.r.

S1.2

Hence

FUNDAMENTAL NOTIONS

23

P(qyr\p) = P(qr\p) + P(q. r\p)+P{'^ q.r\p) (Th. 4),

whence the theorem follows.

It follows that, whether q and r are exclusive or not,

P{q'^r\p) < P(q\p) + P{r\p),

since P(qr |^) cannot be negative. Theorems 4 and 5 together express upper and lower bounds to the possible values of P{q v r | p) irrespective of exclusiveness. It cannot be less than either P{q\p) or P(r|p); it cannot be more than P(q |p) + P(r |p).

Theorem 6. If ^ probable and exclusive

alternatives on data p, and if Q and R are disjunctions of two subsets of these alternatives, of numbers m and n, then P{Q |p)/P(i? |p) min.

For if a is any one of the equal numbers P(?ilp), we

have, by Convention 2,

P{Q Ip) ^ yna\ P{R |p) na\ whence the theorem follows.

Theorem 7. In the conditions of Theorem 6, if q^, exhaustive on data p, and R denotes their disjunction, then R is entailed

by p and P(R\p) -= 1 (Conv. 3).

It follows that P{Q Ip) min.

This is virtually Laplace’s rule, stated at the opening of the Thdorie Analytique. R entails itself and therefore is a possible value of p; hence

P(Q I R) = 7nln.

This may be read: given that a set of alternatives are equally probable, exclusive, and exhaustive, the probability that some one of any subset is true is the ratio of the number in that subset to the whole number of possible cases. This form depends on Convention 3, and must be used only in cases where that convention is adopted. Theorem 6, however, is inde- pendent of Convention 3. If we chose to express certainty on data p by 2 instead of 1 , the only change would be that all numbers associated with probabilities on data p would be multiplied by 2, and Theorem 6 would still hold. Theorem 6 is also consistent with the possibility that the number of alternatives is infinite, since it requires only that Q and R shall be finite subsets. But in this case the number associated with the probability of any infinite subset may be infinite and Convention 3 Js then unsuitable.

24

FUNDAMENTAL NOTIONS

Chap. I

Theorems 6 and 7 tell us how to assess the ratios of probabilities, and, subject to Convention 3, the actual values, provided that the propositions considered can be expressed as finite subsets of equally probable, exclusive, and, for Theorem 7, exhaustive alternatives on the data. Such assessments will always be rational fractions, and may be called i?-probabilities. Now a statement that m and n cannot exceed some given value would be an empirical proposition asserted a priori, and would be inadmissible on rule 5. Hence the i?-probabilities possible within the formal scheme form a set of the ordinal type of the rational fractions.

If all probabilities were /^-probabilities there would be no need for Axiom 5, and the converse of Theorem 2 could hold. But many pro- positions that we shall have to consider are of the form that a magni- tude, capable of a continuous range of values, lies within a specified part of that range, and we may be unable to express them in the required form. Thus there is no need for all probabilities to be im- probabilities. However, if a proposition is not expressible in the required form, it will still be associated with a reasonable degree of belief by Axiom 1, and this, by Axiom 2, Avill separate the degrees for i? -probabilities into two segments, according to the relations 'more probable than' and Tess probable than’. The corresponding numbers, the im-probabilities themselves, will be separated by a unique real number, by Axiom 5 and an application of Dedekind’s section. We take the numerical assessment of the probabilitv of a proposition not expressible in the form required b}^ Theorems 6 and 7 to be this number. Hence we have

Theorem 8. Any probability can be expressed by a real number.

If X is a variable capable of a continuous set of values, we may consider the probability on data p that x is less than x^, say

P{x <Xo\p) =f(Xo).

l{f{xQ) is differentiable we shall then be able to write

P{xo <x < XQ-{-dx^\p) =- f\xQ)dxQ+o (dx^).

We shall usually write this briefly P{dx \p) = f'{x)dx, dx on the left meaning the proposition that x lies in a particular range dx, f\x) is called the probability density.

Theorem 9. If Q is the disjunction of a set of exclusive alternatives on data p, and if R and S are subsets of Q (possibly overlapping) and if

§1.2

FUNDAMENTAL NOTIONS

26

the alternatives in Q are all equally probable on data qi and also 07i data lip, then

P{RS\ p) P(R \p)P(S I Rp)lP(R I Rp).

For suppose that the f)roposition8 contained in Q are of number n, that the subset R contains m of them, and that the part common to R and S contains I of them. Put

P{Q \p) = a.

Then, by Tlieorem 0,

P(R Ij)) inajn] P{RS \p) ■= lajn,

P(S I Rp) is the probability that the true proposition is in the S subset given that it is in the R subset and p, and therefore is equal to {ljw)P(R I Rp). Also RSp entails R; hence

P(S I Rp) P{SR I Rp) (Ax. b)

and

P{RS \p) (ljin){maln) P(R \p)P(S | Rp)IP{R | Rp).

This is the first proposition that we have had that involves probabilities on different data, two of the factors being on data p and two on data Rp. Q itself does not appear in it and is therefore irrelevant. It is introduced into the theorem merely to avoid the use of Convention 3. It might be identical with any finite set that includes both R and S.

The proof has assumed that the alternatives considered are equally probable both on data p and also on data Rp. It has not been found possible to prove the theorem without using this condition. But it is necessary to further developments of the theory that we shall have some way of relating probabilities on different data, and Theorem 9 suggests the simplest general rule that they can follow if there is one at all. We therefore take the more general form as an ax’om, as follows.

Axiom 7. For any q^ropositions p, q, r,

P(qr \p) = P(q \p)P{r \qp)!P(q \qp)-

If we use Convention 3 on data qj) (not necessarily on data p), P(q I qp) r_-: 1 , and we have W. E. Johnson’s form of the product rule, which can be read: the probability of the joint assertion of two propositions on any data p is the product of the qyrobahility of one of them on data p and that of the other on the first and p.

We notice that the probability of the logical sum follows the addition rule (with a caveat), that of the logical product the product rule. This parallel between the Principia and probability language is lost when the joint assertion is called the sum, as has occurred in some recent writings.

26

FUNDAMENTAL NOTIONS

Chap. I

In a sense a probability can be regarded as a logical quotient, since in the conditions of Theorem 7 the probability of Q given R is the probability of Q given p divided by that of R given p. Tin's has been recognized in the history of the notation, which Keynesf traces to H. McColl. McColl wrote Die probability of a, relative to the a priori premiss A, as aje, and relative to bh as a/6. This was modified by W. E. Johnson to ajh and ajhli, and he is followe^d by Keynes, Broad, and Ramsey. Wrinch and I found that this notation was inconvenient when the solidus may have to be used in its usual mathematical sense in the same equation, and introduced P(p\q), which I modified further to P(p\q) in Scientific Inference because the colon was beginning to be needed in the Principia sense of a bracket.

The sum of two classes a an-1 jS, in Principia, is the class y such that every member of ex or of ^ is in y, and conversely. The pn>duct class of (\ and is the class 8 of members common to x and Thus Theorem 5 has a simple analogy with the numbers of members of the classes rx and /5, y and 8. The multiplicative class of lx and ^ is the class of all pairs, one from cx and one from it is this class, n^J the product class, that giv^es an interpretation to the product of the Jiunibers of members of a and jS.

The extension of the product rule from Theorem 9 to Axiom 7 has been taken as axiomatic. This is an application of a principle repeatedly adopted in Principia Mathemalica. If there is a choice between possible axioms, we take the one that enables most consequences to be drawn. Such a generalization is not inductive. What we are doing is to seek for a set of axioms that will permit the construction of a theory of induction, the axioms themselves being primitive postulates. The choice is limited by rule 6; the axioms must be reduced to the minimum number, and the check on whether we make them too general will be provided by rule 2, which will reject a theory if it is found to lead to contradictory consequences. Consider then whether the rule

P{qr \p) == P{q \p)P{r \ qp)

can hold in general. Suppose first that p entails r^ :qr; then either p entails ~ g, or ^ and q together entail ^ r. In either case both sides of the equation vanish and the rule holds. Secondly, suppose that entails qr\ then p entails q and pq entails r. Thus both sides of the equation are 1, Similarly, we have consistency in the converse cases where p

t Treatifte on Probability, 1921, p. 155. This book is full of interesting historical data and contains many important critical remarks. It is not very successful on the con- structive side, since an unwillingness to generalize the axioms has prevented Keynes from obtaining many important results.

FUNDAMENTAL NOTIONS

27

§ 1.2

entails or pq entails ~r, or ^ entails q and pq entails r. This covers the extreme cases.

If there are any cases where the rule is untrue, we shall have to say that in such cases P(qr \p) depends on something besides P(q \ p) and I a new hypothesis would be needed to deal with such cases.

By rule 6, we must not introduce any such hypothesis unless need for it is definitely shown. The product rule may therefore be taken as general unless it can be shown to lead to contradictions. We shall see (p. 35) that consistency can be proved in a wide class of cases.

1.21. The product rule is often misread as follows: the joint proba- bility of two propositions is the product of their probabilities separately. This is meaningless as it stands because the data relative to which the probabilities are considered are not mentioned. In actual application, the rule so stated is liable to become: the joint probability of two pro- positions on given data is the product of their separate probabilities on those data. This is false. We may see this by considering extreme cases. The con'ect statement of the rule may be written (using Convention 3 on data pr)

P(2>q I r) = P(p I r)P(q \pr) (1)

and the other one as

P{pq\r) = P{p\r)P{q\r). (2)

lip cannot be true given r, then p and q cannot both be true, and both (1) and (2) reduce to 0 ~ 0. If^ is certain given r, both reduce to

P{q\r)^ P(q\r) (3)

since in (1) the inclusion of p in the data tells us nothing about q that is not already told us by r. If q is impossible given r, both reduce to 0 = 0. If is certain given r, both reduce to

P(p\r)=^P(p\r), (4)

So far everything is satisfactory. But suppose that q is impossible given pr. Then it is impossible for pq to be true given r, and (1) reduces correctly to 0 = 0. But (2) reduces to

0 = P{p\r)P{q\r),

which is false; it is perfectly possible for both p and q to be consistent with r and pq to be inconsistent with r. Consider the following. Let r consist of the following information: in a given population all the members have eyes of the same colour; half of them have blue eyes and half brown; one member is to be chosen, and any member is equally likely to be selected, p is the proposition that his left eye is blue, q the

28

FUNDAMENTAL NOTIONS

Chap. I

proposition that his right eye is brown. What is the probability, on data r, that his left eye is blue and his right brown ? P(p | r) and P(q \ r) are both J, and according to (2) 1 r) But according

to (1) the probability that his right eye is brown must be assessed subject both to the information that his eyes are of the same colour and that his left eye is blue, and this probability is 0. Thus (1) gives P(pq I r) 0. Clearly the latter result is right; further applications of the former, considering also (left eye brown) and ^q (right eye blue) lead to the astonishing result that on data including the pro- position that all members have two eyes of the same colour, it is as likely as not that any member will have eyes of different colours.

This trivial instance is enough to dispose of (2); but (2) has been widely applied in cases where it gives wrong results, and sometimes seriously wrong ones. The Boltzmann P^-theorem of the kinetic theory of gases rests on a fallacious application of it, since it considers an assembly of molecules, possibly with differences of density from place to place, and gives the joint probability that two molecules will be in adjoining regions as the product of the separate probabilities that they will be. If there are differences of density, and one molecule is in a region chosen at random, that is some evidence that the region is one of high density; then the probability that a second is in the region, given that the first is, is somewhat higher than it would be in the absence of information about the first. Similar considerations apply to Boltz- mann’s treatment of the velocities. In this case the mistake has not prevented the right result from being obtained, though it does not follow from the hypotheses.

Nevertheless there are many cases where (2) is true. If

P{q\pr) = P{q\r)

we say that p is irrelevant to given r.

1.22. Theorem 10. If q^, qn ^ ciUernatives, H the

information already available, and p some additional information, then the ratio

P{q,\pH)P(qMrH)

P(qAH)P{p\q,H)

is the same for all the q^.

By Axiom 7

P(Mr I m = P(P I H)P(q, \pHmp \pH) (1)

= P{q, I H)P{p I q,H)IP(q, \ q,H), (2)

§1.2

FUNDAMENTAL NOTIONS

29

whence

P{q, lpH)P(qr \ qr_H) P(v I pH) P(qAH)P(p\q,H) ~ 'P(p\H)

(3)

which is independent of q^.

Jf we use unity to denote certainty on data q^H for all the (3) becomes

P{qr \pH) oc P(q^ | //)P(p | q^H)

(4)

for variations of q^. This is the principle of inverse probability, first given by Bayes in 1763. It is the chief rule involved in the process of learning from experience. It may also be stated, by means of the product rule, as follows:

P(qr\vB)cc P{pq,\H),

(5)

This is the form used by Laplace, ^ by way of the statement that the posterior probabilities of causes are proportional to the probabilities a priori of obtaining the data by way of those causes. In the form (4), if is a description of a set of observations and the a set of hypotheses, the factor P(g^ | H) may be called the prior probability, P(qj,\pH) the posterior q^'^obability , and P{p\q^H) the likelihood, a convenient term introduced by Professor R. A. Fisher, though in his usage it is sometimes multiplied by a constant factor. It is the proba- bility of the observations given the original information and the hypothesis under discussion. The term a priori probability is sometimes used for the prior probability, but this term has been used in so many senses that the only solution is to abandon it. To Laplace the a priori probability meant P(pqr \H), and sometimes the term has even been used for the likelihood. A priori has a definite meaning in logic, in relation to propositions independent of experience, and we frequently have need to use it in this sense. We may then state the principle of inverse probability in the form: The posterior 'probabilities of the hypo- theses are proportional to the products of the prior probabilities and the likelihoods. The constant factor will usually be fixed by the condition that one of the propositions g^ to g,^ must be true, and the posterior probabilities must therefore add up to 1. (If 1 is not suitable to denote certainty on data pH, no finite set of alternatives will contain a finite fraction of the probability. The rule covers all cases when there is anything to say.)

The use of the principle is easily seen in general terms. If there is originally no ground to believe one of a set of alternatives rather than another, the prior probabilities are equal. The most probable, when evidence is available, will then be the one that was most likely to lead to that evidence. We shall be most ready to accept the hypothesis that

30

FUNDAMENTAL NOTIONS

Chap. I

requires the fact that the observations have occurred to be the least remarkable coincidence. On the other hand, if the data were equally likely to occur on any of the hypotheses, they tell us nothing new with respect to their credibility, and w^e shall retain our previous opinion, whatever it was. The principle wdll deal with more complicated circum- stances also; the immediate point is that it does provide us with what we want, a formal rule in general accordance with common sense, that will guide us in our use of experience to decide betw^een hypotheses.

1.23. We have not yet showm that Convention 2 is a convention and not a postulate. This must be done by considering other possible conven- tions and seeing what results they lead to. Any other convention must not contradict Axiom 4. For instance, if the number associated with a probability by our rules is x, we might agree instead to use the number Then if x and x' are the present estimates for the propositions q and q\ and for r and r', those for q^q' and rvr' will both be and the consistency rule of Axiom 4 will be satisfied. But instead of the addition rule for the number to be associated with a disjunction we shall have a product rule. Every proposition stated in either notation can be translated into the other; if our present system leads to the result that a hypothesis is as likely to be true as it is that we should pick a white ball at random out of a bag containing 99 white ones and 1 black one, that result will also be obtained on the suggested alternative system. The fundamental notion is that of the comparison of reasonable degrees of belief, and so long as all methods place them in the same order the differences between the methods are conventional. This will be satisfied if instead of the number x we choose any function of it,f(x), such that X Eindf(x) are increasing functions of each other, so that for any value of one the other is determinate. This is necessary by Convention 1 and Axiom 1, but every form off{x) will lead to a different rule for the probability -number of a disjunction if it is to be consistent with Axiom 4. Hence the addition rule is a convention. It is, of course, much the easiest convention to use. To abandon Convention 1, con- sistently with Axiom 1 , would merely arrange all numerical assessments in the opposite order, and again the same results would be obtained in translation. The assessment by numbers is simply a choice of the most convenient language for our purposes.

1.3. The original development of the theory, by Bayes,! proceeds differently. The foregoing account is entirely in terms of rules for the

t Phil. Tram. 53, 1763, 376-98.

§1.3 FUNDAMENTAL NOTIONS 31

comparison of reasonable degrees of belief. Bayes, however, takes as his fundamental idea that of expectation of benefit. This is partly a matter of what we want, which is a separate problem from that of what it is reasonable to believe; I have therefore thought it best to proceed as far as })ossible in terms of the latter alone. Nevertheless, we have in practice often to make decisions that involve not only belief but the desirability of the possible effect of different courses of action. If we have to give advice to a practical man, either we or he must take these into account. In deciding on his course of action he must allow both for the probability that the action chosen will lead to a certain result and for the value to him of that result if it happens. The fullest development on these lines is that of F. P. Ramsey. | I shall not attempt to reproduce it, but shall try to indicate some of the principal points as they occur in his work or in Bayes’s. The fundamental idea is that the values of expectations of benefit can be arranged in an order; it is legitimate to compare a small probability of a large gain with a large probability of a small gain. The idea is necessarily more com- plicated than my Axiom 1; on the other hand, the comparison is one that a business man often has to make, whether he wants to or not, or whether it is legitimate or not. The rule simply says that in given circumstances there is always a best way to act. The comparison of probabilities follows at once; if the benefits are the same, wliichever of two events happens, then if the values to us of the expectations of benefit differ it is because the events are not equally likely to happen, and the larger value is associated with the larger probability. Now we have to consider the combination of expectations. Here Bayes, I think, overlooks the distinction between what Laplace calls ‘mathematical’ and ‘moral’ expectation. Bayes spealis in terms of monetary stakes, and would say that a 1/100 chance of receiving £100 is as valuable as a certainty of receiving £1 . A gambler might say that it is more valuable; most people would perhaps say that it is less so. Indeed Bayes’s definition of a probabilit}^ of 1/100 would be that it is the probability such that the value of the chance of receiving £100 is the same as the value of a certain £1. Since different values may be compared, the uniqueness of a probability so defined requires a postulate that the value of the expectation, the proposition and the data remaining the same, is proportional to the value to be received if the proposition

t The Foundations of MathenuUics, 1931. pp. 157-211. This essay, like tliat of Bayes, was published after tiie author’s death, and suffers from a number of imperfections in the v^erbal statement that he might have corrected.

32

FUNDAMENTAL NOTIONS

Chap. I

is true. This is taken for granted by Bayes, and Ramsey makes an equivalent statement (foot of p. 179). The difficulty is that the value of £l to us depends on how much money we have already. This point was brought out by Daniel Bernoulli in relation to what was called the Petersburg Problem. Two players play according to the following rules. A coin is to be thrown until a head is thrown. If it gives a liead on the first throw, J is to pay £ £1 ; if the first head is on the second throw, £2; on the third, £4; and so on. What is the fair sum for B to pay A for his chances? The mathematical expectation in ])ounds is

'2 1 “f" J 2 "P "8 OC .

Thus on this analysis B should pay A an infinite sum. If we merely consider a large finite sum, such as £2^^\ he will lose if there is a head in any of the first 20 throws; he will gain considerably if tiie first liead is on the 21st or a later throw. The question was, is it really worth anybody’s while to risk such a sum, most of wliicli he is practically certain to lose, for an almost inappreciable chance of an enormous gain? Even eighteenth -century gamblers seem to have had doubts about it. Daniel Bernoulli’s solution was that tltC value of £22^‘ is very different according to the amount we have to start with. The value of a loss of that sum to anybody that has just that amount is not equal and opposite to the value of a gain of the same sum. He suggested a law relating the value of a gain to the amount already possessed, which need not detain us;t but the impprtant point is that he recognized that expectations of benefit are not necessarily additive. What Laplace calls ‘moral expectation’ is the value or pleasure to us of an event; its rela- tion to the monetary value in terms of mathematical expectation may be rather remote. Bayes wrote after Bernoulli, but before Laplace, but he does not mention Bernoulli. Nevertheless, the distinction does not dispose of the interest of the treatment in terms of expectation of benefit. Though we cannot regard the benefits of gains of the same kind as mutually irrelevant, on account of this psychological pheno- menon of satiety, there do seem to be many cases where benefits are mutually irrelevant. For instance, the pleasures to me of two dinners on consecutive nights seem to be nearly independent, though those of two dinners on the same night are definitely not. The pleasures of the unexpected return of a loan, having a paper accepted for publication, a swim in the afternoon, and a theatre in the evening do seem

t It is that the value of a gain dx, when we liavo x already, is proportional to dx/x; this is the rule associated in certain biological applications with the names of Weber and Fechner.

FUNDAMENTAL NOTIONS

33

§ 1.3

independent. If tliere are a sufficient number of such benefits (or if there could be in some possible world, since all we need is consistency), a scale of the values of benefits can be construct(;d, which will satisfy the commutative rule of addition, and then, by Bayes’s principles, one of probability in terms of them. The addition rule will then be a theorem. The product rule is treated by Bayes in the following v ay. We can write E(a,]) \q) for the value of the expectation of receiving a if p is true, given q, and by definition of F(p | q),

E{a,p\q) aP{p\q).

The proportionality of E{a,p \ q) to a, given p and q, is a postulate, as we have already stated. Consider the value of the expectation of getting a if p and q are both true, given r. This is (iP(pq | r). But we may test 2^ ^^d then q. If p turns out to be true, our expectation

will be (iP(q \ pT), since jy is now among our data; if untrue, we know that we shall receive nothing. Now return to the first stage. If p is true we shall receive an expectation, whose value is aP(q | pr), otherwise nothing. Hence our initial expectation is aP{q | pr)P{p | r); whence

P{pq I r) =- P(p 1 r)P\q | pr).

Ramsey’s presentation is much more elaborate, but depends on the same main ideas. The proof of the principle of inverse probability is simple. The difficulty about the separation of propositions into dis- junctions of equally possible and exclusive alternatives is avoided by this treatment, but is replaced by difficulties concerning additive expec- tations. These are hardly practical ones in either case; no practical man will refuse to decide on a course of action merely because we are not quite sure which is the best v ay to lay the foundations of the theory. He assumes that the course of action that he actually chooses is the bestji Bayes and Ramsey merely make the less drastic assumption that there is some course of action that is the best. In my method expectation would be defined in terms of value and probability; in theirs probability is defined in terms of values and expectations. The actual propositions are of course identical.

1.4. At any stage of knowledge it is legitimate to ask about a given hypothesis that is accepted, ‘How do you know?’ The answer will usually rest on some observational data. If we ask further, ‘What did you think of the hypothesis before you had these data? we may be told of some less convincing data; but if we go far enough back w^e shall always reach a stage where the answer must be : M thought the matter

3595.58 T)

34

FUNDAMENTAL NOTIONS

Chap. I

worth considering, but had no opinion about whether it was true.’ What was the probability at this stage ? We have the answer already. If there is no reason to believe one hypothesis rather than another, the probabilities are equal. In terms of our fundamental notions of the nature of inductive inference, to say that the probabilities are equal is a precise way of saying that we hove no ground for choosing between the alternatives. All hypotheses that are sufficiently definitely stated to give any difference between the probabilities of their consequences will be compared with the data by the principle of inverse probability; but if we do not take the prior probabilities equal we are expressing confidence in one rather than another before the data are available, and this must be done only from definite reason. To take the prior probabili- ties different in the absence of observational reason for doing so w ould be an expression of sheer prejudice. The rule that we should then take them equal is not a statement of any belief about the actual composi- tion of the world, nor is it an inference from previous experience; it is merely the formal way of expressing ignorance. It is sometimes referred to as the Principle of Insufficient Reason (Laplace) or the equal dis- tribution of ignorance. Bayes, in his great memoir, repeatedly says that the principle is to be used only in cases w here we have no grt>und whatever for choosing between the alternatives. It is not a new rule in the present theory because it is an immediate application of (Vjnven- tion 1. Much confusion has arisen about it through misunderstanding and attempts to reinterpret it in terms of frequency definitions. My contention is that tlie frequency definitions themselves lead to no results of the kind that we need until the notion of reasonable degree of belief is reintroduced, and that since the whole purpose of these definitions is to avoid this notion they necessarily fail in their object. When reasonable degree of belief is taken as the fundamental notion the rule is immediate. We begin by making no assumption that one alternative is more likely than another and use our data to compare them .

Suppose that one hypothesis is suggested by one person A, and another by a dozen B, C,...; does that make any difference? No; but it means that we have to attend to tw^o questions instead of one. First, is p or g true ? Secondly, is the difference between the suggestions due to some psychological difference between A and the rest? The mere voting is not evidence because it is quite possible for a large number of people to make the same mistake. The second question cannot be answered until we have answered the first, and the first must be con- sidered on its merits apart from the second.

FUNDAMENTAL NOTIONS

35

§1.5

1.5. We are now in a position to consider whether we have fulfilled the conditions that we required at the outset. I think ( 1 ) is satisfied, though the history of both probability and deductive logic is a warning against over 'Confidence that an unstated axiom has not slipped in.

2. Axiom 1 assumes consistency, but this assumption by itself does not guarantee that a given system is consistent. It makes it possible to derive theorems by equating probabilities found in different ways, and if in spite of all efforts probabilities found in different ways were different, the axiom would make it impossible to accept the situation as satisfactory. We must not expect too much in the nature of a general proof of consistency. There is a theorem due to G5del that if any logical system that includes arithmetic contained a proof of its own consis- tency. it would also contain one of its own inconsistency; so apparently it would be fatal to a system if we could find a general proof of consis- tency within it. Proofs of the consistency of various logical schemes (including the system of Principla Matheniatica and therefore the theory of functions of a real variable) do exist, but only by going out- side the frames of the schemes themselves. The proof amounts to finding a proposition that can be stated in the system but cannot be proved or disproved by using the rules of the system. Since the system of Principia contains a proposition that two contradictory propositions imply any proposition, the existence of an undemonstrable proposition implies that the primitive propositions in the system are consistent. But this argument itself cannot be expressed in Principia language! What we want is that the probability of a proposition on the same data shall always be the same; thus, if w^e are considering two alternative hypotheses and our previous information is //, and the new evidence consists of two batches of data p^ and p^, the assessments on data pj^p^^H should be the same whether we take p^ or pg account first or both at once. Now, by the principle of inverse probability,

P(<li \PiH) ^ P(q^\p^H)

Phi\H)P{Pi\<liP) Pi<iAP)P{Pi\%Py

Replacing H hy p^H and p^ by p^ we shall obtain the result for the application of the additional data p^, Pi being now already given :

P(9[i \PiP%H) Pj^ilPxPzH)

P(<i\ \PiP)P(P2\qiPiP) P(qi\PiP)P{P2\qiPiii)'

Multiplying, we have

P{q^ IPiP^H) ^ PjqtlPiPzH)

P{qi\P)P(Pi\qiH)P(P2 \PiqiH) P(qi\H)P(Pi\q^H)P(p^ \PiqtH)'

36 FUNDAMENTAL NOTIONS Chap. I

But by Axiom 7, assuming that the product rule holds for likelihoods, P(Pi\q\H)P(Pi\PiqiH) =. P(PxPt\q-^B).

and therefore

^(gi \V\PzP) ^ Pili \P\P-iP) _

P{9i \P)PiPiPz\QiP) P(9z\P)P(PiP2 1^2^)’ which is the result of applying the principle of inverse probability to take account of the data and p^ simultaneously. By symmetry we should obtain the same result if we took account of pg fir«t. Extension to any number of batches of new data is obviously possible, and the results will therefore be consistent provided that we always start with the same data and finish wdth the same, and that w^e take account of the new' data as we proceed. Keglect of the last condition may lead to inconsistencies, but that is the result of not applying the principle correctly. In the proof we have assumed that the product rule holds for likelihoods. This has not been proved in general, but has invariably been assumed even by those who claim to reject the principle of inverse probability. What our theorem shows is that if the product rule holds for likelihoods the principle of inverse probability cannot lead to contradiction.

The consistency of the product rule can be treated more diiectly as follows. Let be two sets of propositions each exclusive and

exhaustive on p, and denote their disjunctions by Q, R. Then

Pifk lp) = P{Qrk lg>) = I PiqiTk \p)-

i

Instead of Axiom 7 assume that

P(ii\^k'P) P^j^klp)'

and assume that probabilities on data p satisfy the axioms. Then for probabilities on data it is obvious that Axioms 1, 2, 5 are satisfied; Axioms 3, 4, 6, 7 are easily proved, beginning with Axiom 6. Hence if we weaken Axiom I to a statement that probabilities are comparable given one sufficiently wide datum p, we can consistently convert the product rule into a definition of probabilities on data including p.

3. For any assessment of the prior probability the principle of inverse probability will give a unique posterior probability. This can be used as the prior probability in taking account of a further set of data, and the theory can therefore always take account of new information. The choice of the prior probability at the outset, that is, before taking into account any observational information at all, requires further con- sideration. We shall see that further principles are available as a guide.

§1.6

FUNDAMENTAL NOTIONS

37

These principles sometimes indicate a unique choice, but in many problems some latitude is permissible, so far as we know at present. In such cases, and in a different world, the matter would be one for decision by the International Research Council. Meanwhile we need only remark that the choice in practice, within the range permitted, makes very little difference to the results.

4. This is satisfied by definition.

5. We have avoided contradicting rule 5 so far, but further applica- tions of it will appear later.

6. Our main postulates are the existence of unique reasonable degrees of belief, which can be put in a definite order; Axiom 4 for the consistency of probabilities of disjunctions; either the axiomatic extension of the product rule or the theory of expectation. It does not appear that these can be reduced in number, without making the theory incapable of covering the ground required.

7. The simple cases mentioned on pp. 29-30 show how the principle of inverse probability does correspond to ordinary processes of learning, though we shall go into much more detail as we proceed. Differences between individual assessments that do not agree with the results of the theory will be part of the subject-matter of psychology. Their existence can be admitted without reducing the importance of a unique standard of reference. It has been said that the theory of probability could be accepted only if there was experimental evidence to support it; that psychology should invent methods of measuring actual degrees of belief and compare them wuth the theory. I should reply that without an impersonal method of analysing observations and drawing inferences from them we should not be in a position to interpret these observations either. The same considerations would apply to arithmetic. To quote P. E. B. Jourdainit

T Hornetirnes feel inclined to apply the historical method to the multiplication table. I should make a statistical inquiry among school children, before their pristine wisdom had been biased by teachers. I should put down their answers as to what 6 times 9 amounts to, I should work out the average of their answers to six places of decimals, and should then decide that, at the present stage of human development, this average is the value of 6 times 9.’

I would add only that without the multiplication table we should not be able to say what the average is. Nobody says that wrong answers invalidate arithmetic, and accordingly we need not say that the fact that some inferences do not agree with the theory of probability

t T/ie Philosophy of Mr, B*rtr*nd R* 88*11, 1918, p. 88.

38

FUNDAMENTAL NOTIONS

Chap. I

invalidates the theory. It is sufficiently clear that the theory does represent the main features of ordinary thought. The advantage of a formal statement is that it makes it easier to see in any particular case whether the ordinary rules are being followed.

This distinction shows that theoretically a probability should always be worked out completely. We have again an illustration from pure mathematics. What is the 1,000th figure in the expansion of Nobody knows; but that does not say that the probability that it is a 5 is 0- 1 . By following the rules of pure mathematics we could deter- mine it definitely, and the statement is either entailed by the rules or contradicted; in probability language, on the data of pure mathematics it is either a certainty or an impossibility.! Similarly, a guess is not a probability. Ih'obability theory is more complicated than deductive logic, and even in pure mathematics we must often be content with approximations. Mathematical tables consist entirely of approxima- tions. Bence we must expect that our numerical estimates of proba- bilities in practice will usually be approximate. The theory is in fact the system of thought of an ideal man that entered the world knowing nothing, and always worked out his inferences completely, just as pure mathematics is part of the system of thought of an ideal man who always'gets his arithmetic right. J But that is no reason w hy the actual man should not do his best to approximate to it.

1.6. We can now^ indicate in general terms how^ an inductive inference can approach certainty, though it cannot reach it. If g is a hypothesis, H the previous information, and an experimental fact, we have by tw o applications of the product rule, using Convention 3,

' P(P,\H)

(1)

since both are equal to Pip^q | H^Pip^ \ H). If p, is a consequence of q, P(Pi I qH) ~ 1 ; hence in this case

P(q\pJJ)=.

PjglH)

PiPiim'

(2)

t It is unfortunate that pure mathematicians speak of, for instance, the probability distribution of prime numbers, meaning a smoothed density distribution. Systematic botanists and zoologists are far ahead of mathematicians and physicists in tidying up their language. ® ^

I An expert computer does not trust his arithmetic without applying checks, which would give identities if the work is correct but would be expected to fail if there is a mistake. Thus induction is used to check the correctness of what is meant to be deduction. The possibility that two mistakes have cancelled is treated as so improbable that it can be ignored.

FUNDAMENTAL NOTIONS

39

§ 1.6

If Pz '" further consequences of q, which are found to be true, we shall have in succession

P{<l\PiPzH) ^ P(q I VlP2-‘PnH)

P(q\H)

P(P^\H)P(p^\p,Hy *•*’

P(Vl \H)P{P2\Plii)- P\Pn \ Vv-Pn~lPy

(3)

Thus each verification divides the probability of the hypothesis by the probability of the verification, given the previous information. Thus, with a sufficient number of verifications, one of three tilings must happen: (1) The probability of q on the information available will exceed 1. (2) it is always 0. (3) P{p,, \ jhP2'-Vn-iP) 1-

(1) is impossible since the highest degree of probability is certainty.

(2) means that q can never reach a finite probability, however often it is verified. But if w^e adopt (3), repeated verifications of consequences of a hypothesis w ill make it practically^ certain that further consequences of it will be verified. This accounts for the confidence that w e actually have in inductive inferences.

This proposition also provides us with an answer to various logical difficulties connected w ith the fact that if p entails </, q does not neces- sarily entail p. p may be one of many alternatives that would also entail q. Tn the lowest terms, if q is the disjunction of a set of alterna- tives ^1, qw fhen any member of this set entails q, but q does not

entail any particular member. Now in science one of our troubles is that the alternatives available for consideration are not ahvays an exhaustive set. An unconsidered one may escape attention for centuries. The last proposition shows that this is of minor importance. It says that if pj,..., p^^ are successive verifications of a hypothesis q,

P{Vn \PlP2 -Pn-xP)

will approach certainty; it does not involve q and therefore holds whether q i.s true or not. The unconsidered hypothesis, if it had been thought of, would either (1) have led to the consequences pj, pgv or (2) to different consequences at some stage. In the latter case the data would have been enough to dispose of it, and the fact that it was not thought of has done no harm. In the former case the considered and the unconsidered alternatives would have the same consequences, and will presumably continue to have the same consequences. The un- considered alternative becomes important only when it is explicitly stated and a type of observation can be found where it would lead to different predictions from the old one. The rise into importance of the

40

FUNDAMENTAL NOTIONS

Chap. I

theory of general relativity is a case in point. Even though we now know that the systems of Euclid and Newton need modification, it was still legitimate to base inferences on them until we knew what particular modification was needed. The theory of probability makes it possible to respect the great men on whose shoulders we stand.

Tlie possibility of this procedure rests, of course, on the fact that there are cases where a large number of observations have been found to agree with predictions made by a law. The interest of an estimate of the probability of a law, given certain data, is not great unless tliose actually are our data. Indeed, a statement of it might lead to highly uncomplimentary remarks. It is not necessary that the predictions shall be exact. Jn the case of uniformly accelerated motion mentioned near the beginning, if the law is stated in the form that at any instant tj. the observed .s will lie between where e is small com-

pared with the whole range of variation of .v, it w ill still be a legitimate inference after many verifications that the law will hold in future instances within this margin of uncertainty. This takes us a further step towards understanding the nature of the acceptance of a simple law^ in spite of the fact that in the criiue form given in applied mathe- matics it does not exactly agree with the observations.

1.61. If we lump together all hypotheses that give indistinguishable consequences, their total probability will tend to 1 with sufiicient verification. For if we have a set of hypotheses r/p..., all asserting that a quantity x will lie in a range ±e, we may denote their disjunction by q, w^hich will assert the same. Suppose that q would permit the quantity to lie in a range ± E, w here E is much greater than €. Suppose further that x is measured and found to be in the range indicated by q. Then if p denotes this proposition, P{j) \qh) ~ 1, and P{p | ^ qli) is of order ejE. Hence

P(q\ph) _ (E\ P(q\h)

P(r^q\ph) \e/P(~g'|A)'

Thus if E/e is large and ^ is a serious possibility, a single v^erification ma}^ send its probability nearly up to 1 . It is an advantage to consider together in this way all hypotheses that would give similar inferences and treat their disjunction as one hypothesis. The data give no informa- tion to discriminate between them vso long as the data are consequences of all; the posterior probabilities remain in the ratios of the prior probabilities. With this rule, therefore, we can with a few verifications exclude from serious consideration any vaguely stated hypotheses that would require the observed results to be remarkable coincidences; while

§1.6

FUNDAMENTAL NOTIONS

41

unforeseen alternatives whose consequences would agree with those given by hypotheses already included in g, within the range of verifica- tion at any stage, will give no trouble. By the time when any of them is stated explicitly, all hypotheses not implying values of x within the ranges actually found will have negligible probabilities anyhow, and all that we shall need to do is to separate the disjunction q as occasion arises. It is therefore desirable as far as possible to state hypotheses in such a form that those with indistinguishable consequences can be treated togctlier; this will avoid mere mathematical complications relating to possibilities that we have no means of testing.

1.7. I'jiEORKM II. // alferna lives on

data r, and if

- - -- A?' I '?„»•).

the7i each ^ P{p \q^ v y,,... ^ q„ \r).

For if we denote the disjunction v rp,... v q^^ by q^ we have

P{pq j /•) P{pq^ I r)-t P{pq.^ \ r)^ ... (1 )

since these alternatives are mutually exclusive'; and this

- (2) The first factors are all etpial, and the sum of the second factors is P{q I r). Hence

P(pq I r) P(p I q^ r)P(q \r). (3)

But P(pq I r) ^ P{p 1 qr)P{q | r), (4)

which gives the theorem on comparison witli (3).

This leads to the principle that we may call the suppression of an irrelevant premiss. If (7i entailed by r,

P{p i qr) =- P(2)q | r) P(p | r),

since P(q \pr) ^ 1; and then each of the expressions P{p |g^r) is equal to P{p I r). In words, it the probability of a proposition is the same for all the alternative data consistent with one fixed datum, then the probability on the fixed datum alone has the same value.

The interest of this theorem is primarily in relation to what are called ‘chances’, in a technical sense given by N. R. Campbell and M. S. Bartlett. We have seen that probabilities of propositions in general depend on the data. But cases can be stated, and whether they exist or not must be considered, where the probability is the same over a wide range of data; in such a case we may speak of the information not common to all these data as irrelevant to the probability of the

42 FUNDAMENTAL NOTIONS Chap. I

proposition. Thus above we can say that the propositions irrelevant to j), given r. Further,

P(ni I r) - P{qi I r)P(j) i q^r) -- Pfe | r)P{p | r),

so that the product formula in such a case is legitimately replaced by the form (2) on p. 27. I shall therefore define a chance as follows: If

q,^ are a set of alternatives, miitimlly exclusive and exhaustive on

data r, and if the probabilities of p qitwn any of them and r are the same, each of these probabilities is called the chance of p on data r. It is equalf to P{p I r).

In any case where r includes the specification of all the parameters in a law, and the results of previous trials are irrelevant to the result of a new trial, the probability of a given result at that trial is the chance on data r. For the information available just before that trial is made is composed of r and the results of all previous trials. If we consider the aggregate of all the results that might have been obtained in pre- vious trials, they constitute a set of alternatives such that one of them must occur on data r, and are exclusive and exhaustive, (dven then that the probability of an event at the next trial is the same whatever the results of previous trials, it must be equal to the chance on data r. It follows that the joint probability on data r of the results of several trials is the product of their separate chances on data r. This can easily be proved directly. For if --’Pm results in order, we have

by successive applications of the product formula

k) = P{Vx\'<')P{V% \Pir)P(P3 \VlIh^)-P(Vm\PvPm-l1^)^ and by the condition of irrelevance this is equal to P{Pi I 'r)P(Pi I r)P(p^ I r)...P{p^ I r).

This is usually taken for granted, but it is just as well to have it proved.

When the probabilities, given the law, are chances, they satisfy the product rule automatically. Hence our proof of the consistency of the principle of inverse probability is complete in all cases where the likeli- hoods are derived from chances. This covers nearly all the applications in this book.

Theorem 12. 5'n of alternatives,

each exclusive and exhaustive on data r, and if

P{P,qt\^) = f(P,)9(it)

I Bayes and Laplace use both words ‘probability’ and ‘chance’, but so far as I know do not specify any distinction between them. There are, however, passages in their writings that suggest that they use the words with their modern senses interchanged.

FUNDAMENTAL NOTIONS

43

S 1.7

for all values of s and t, where f(pf\ depends only on p^ and r, and g(q^ only on qi and r, then

PiPs \ »•) oc f(p,)\ P(qt I r) cc g(qi).

For if we denote the disjunctions of the and q( by p and q, we have

P{P. P{Ps q,\r)= f(pf) ^ ( 1 )

t t

which is proportional to f(pg). But

P{Ps<l 1 r) = P(p^ 1 r)P(q \p^r) (2)

and the last factor is 1 since q is entailed by r. Hence

P(Ps\r)^f{pfl- (3)

Similarly, P(q^ [ r) oc g(qi). (4)

We notice that

Pipg. I »•) = 2 I,fiPs)9{q,) = J,f(Ps) 1 9(9t) (6)

8 ( St

and is equal to 1 since p and q are both entailed by r. It is possible to multiply /(p^) and g{qi) by factors such that both sums will be equal to 1 ; these factors will be reciprocals; and if this is done, since p and q separately are entailed by r, we shall have

P(Ps I r) = /(pj; P(qi 1 r) = giq,).

Also P(Pg I g,r) = P(p^qi \ r)IP(qi \ r) = f(p,) (6)

and q^ is irrelevant to p^.

This theorem is useful in cases where a joint probability distribution breaks up into factors.

1.8. Expectation of benefit is taken as a primitive idea in the Bayes- Ramsey theory. In the present one we can define the expectation of a function f(z) on data p by the equation

E{f{x)\p} = 2f(^)P(x\p)

taken over all values of ;r. For expectation of benefit, if benefits inter- fere, there is no great trouble. If x is, for instance, a monetary gain, we need only distinguish between x itself, the expectation of which will be Y,xP{x \ p), and the benefit to us of x, which is not necessarily proportional to x. If it is f(x), the expectation of benefit will be

2/(x) p(x\p).

The expectations of functions of a variable are often required for our purposes, though we shall not have much more to say about expecta- tion of benefit. But attention must be called at once to the fact that if the expectation of a variable is a, it does not mean that we expect

44

FUNDAMENTAL NOTIONS

Chap. I

the variable to be near a. (bnsider the following case. Suppose that we have two boxes A and B each containing n balls. We are to toss a coin; if it comes down heads we shall transfer all the balls from A to B\ if tails, all from B to A. What is our present ex])ectation of the number of balls in A after the process? There is a probability 1 that there will be "In balls in A, and a probability I that there will be none. Hence the expectation is n, which is not a possible value at all. Incor- rect results have often been obtained by taking an expectation as a prediction of an actual value; this can be done only if it is also shown that the probabilities of dilferent actual values are closely concentrated about the expectation. It may easily happen that they are concen- trated about two or more values, none of which is anywhere near the expectation.

1.9. It may be noticed that the words ‘idealism’ and ‘realism’ have not yet been used. I should perhaps explain tiiat their use in everyday speech is different from the philosophical use. In everyday use, realism is thinking that other people are worse than they are; idealism is thinking that they are better than they are. The former is an expres- sion of praise, the latter of disparagement. It is recognized that nobody sees himself as othei’s see him; it follows tliat everybody knows that everybody else is either a realist or an idealist. In pliilosopliy, realism is the belief that there is an external world, which would still exist if we were not available to mak(' observations, and that the function of scientific method is to find out properties of this world. Idealism is the belief that nothing exists but the mind of the observer or observers and that the external world is merely a mental construct, imagined to give us ourselves a convenient way of describing our experiences. The extreme form of idealism is solipsism, which, for any individual, asserts that only his mind and his sensations exist, other people’s minds also being inventions of his own. The methods developed in this book are consistent with some forms of both realism and idealism, but not with solipsism; they contribute nothing to the settlement of the main ques- tion of idealism versus realism, but they do lead to the rejection of various special cases of both. I am personally a realist (in the philo- sophical sense, of course) and shall speak mostly in the language of realism, which is also the language of most people; but if an idealist wishes to translate anything in this book into the language of idealism, I think he will be able to do it. To him I offer the bargain of the Unicorn with Alice: ‘If you’ll believe in me, I’ll believe in you.’

§1.9

FUNDAMENTAL NOTIONS

46

Solipsism is not, as far as I know, actively advocated by anybody (with the possible exception of the behaviourist psychologists). The great difficulty about it is that no two solipsists could agree. If A and B are solipsists, A thinks that he has invented B and vice versa. The relation between them is that between Alice and the Red King; but while Alice Avas willing to believe that she was imagining the King, she found the idea that the King was imagining her quite intolerable. Tweedledum and Tweedledce solved the problem by accepting the King’s solution and rejecting Alice’s; but every solipsist must have his own separate solipsism, which is flatly contradictory to every other’s. Nevertheless, solipsism does contain an important principle, recognized by Karl I^earson, that any person’s data consist of his own individual experiences and that his opinions are the result of his own individual thought in relation to those experiences. Any form of realism that- denies this is sim])ly false. A hypothesis does not exist till some one ])erson has thought of it; an inference does not exist until one person has made it. We must and do, in fact, begin with the individual. But early in life he recognizes groups of sensations that habitually occur together, and in particular he notices resemblances between those groups that we, as adults, call observations of oneself and other people. When he learns to speak he has already made the observation that some sounds belonging to i-hese grouy)s are habitually associated with other groups of visual or tactile sensations, and has inferred the rule that we should express by saying that particular things and actions are denoted by particular words; and when he himself uses language he has generalized the rule to say that it may be expected to hold for future events.

Thus the use of language depends on the principle that generalization from experience is possible; and this is far from being the only such generalization made in infancy. But if we accept it in one case we have no ground for denying it in another. But a person also observes similarities of appearance and behaviour between himself and other people, and as he himself is associated with a conscious personality, it is a natural generalization to suppose that other people are too. Thus the departure from solipsism is made possible by admitting the pos- sibility of generalization. It is now possible for two people to under- stand and agree with each other simultaneously, which would be impossible for two solipsists. But we need not say that nothing is to be believed until everybody believes it. The situation is that one person makes an observation or an inference; this is an individual act. If he

46

FUNDAMENTAL NOTIONS

Chap. I

reports it to anybody else, the second person must himself make an individual act of acceptance or rejection. All that the iirst can say is that, from the observed similarities between himself and other people, he would expect the second to accept it. The facts that organized society is possible and that scientific disagreements tend to disappear when the participants exchange their data or when new data accumu- late are confirmation of this generalization. Regarded in this way the resemblance between individuals is a legitimate induction, and to take universal agreement as a primary requisite for belief is a superfluous postulate.

Whether one is a realist or an idealist, the problem of inferring future sensations arises, and a theory of induction is needed. Both some realists and some idealists deny this, holding that in some way future sensations can be inferred deductively from some intuitive knowledge of the possible properties of the world or of sensations. If experience plays any part at all it is merely to fill in a few details. This must be rejected under rule 5. I shall use the adjective ‘naive for any theory, whether realist or idealist, that maintains that inferences beyond the original data are made with certainty, and ‘critical’ for one that admits that they are not, but nevertheless have validity. Nobody that ever changes his mind through evidence or argument is a naive realist, though in some discussions it seems to be thought that there is no other kind of realism. It is perfectly possible to believe that we are finding out properties of the world without believing that anything we say is necessarily the last word on the matter.

It should be remarked that some philosophers define ‘naif realism’ in some such terms as the belief that the external world is something like our perception of it’, and argue in its favour. To quote a remark I once heard Russell make, ‘I wonder what it feels like to think that.’ The succession of two-dimensional impressions that we call visual observations is nothing like the three-dimensional world of science, and I cannot think that such a hypothesis merits serious discussion. The trouble is that many philosophers are as far as most scientists from appreciating the long chain of inference that connects observation with the simplest notions of objects, and many of the problems that take up most attention are either solved at once or are seen to be insoluble when we analyse the process of induction itself.

II

DIRECT PROBABILITIES

‘Having thus exposod tho far-seeing Mandarin’s inner thoughts, would it be too excessive a labour to jxjiietrate a little deeper into the rich mine of strat^egy and disclose a specific detail ?

Ernest Bramah, Kai Lung Unrolls his Mat

2.0. We have seen that the principle of inverse probability can be stated in the form

Posterior Probability oc Prior Probability x Likelihood,

where by the likelihood we understand the probability that the observa- tions should have occurred, given the hypothesis and the previous knowledge. The prior probability of the hypothesis has nothing to do with the observations immediately under discussion, though it may depend on previous observations. Consequently the whole of the in- formation contained in the observations that is relevant to the posterior probabilities of different hypotheses is summed up in the values that they give to the likehhood. In addition, if the observations are to tell us much that we do not know already, the likelihood will have to vary much more between different hypotheses than the prior probability does. Special attention is therefore needed to the discussion of the probabilities of sets of observations given the hypotheses.

Another consideration is that we may be interested in the likelihood as such. There are many problems, such as those of games of chance, where the hypothesis is trusted to such an extent that the amount of observational material that would induce us to modify it would be far larger than will be available in any actual trial. But we may want to predict the result of such a game; or a bridge player may be interested in such a problem as whether, given that he and his partner have nine trumps between them, the remaining four are divided two and two. This is a pure matter of inference from the hypothesis to the probabili- ties of different events. Such problems have already been treated at great length, and I shall have little to say about them here, beyond indicating their general position in the theory.

In Chapter I we were concerned mainly with the general rules that a consistent theory of induction must follow. They say nothing about what laws actually connect observations; they do provide means of choosing between possible laws, in accordance with their probabilities

48

DIRECT PROBABILITIES

Chap. II

given the observations. The laws themselves must be suggested before they can be considereci in terms of the rules and the observations. The suggestion is always a matter of imagination or intuition, and no general rules can be given for it. We do not assert that any suggested hypo- thesis is right, or that it is wrong; it may appear that there are cases where only one is available, but any hypothesis specific enougli to give inferences has at least one contradictory, in comparison with which it may be considered. The evaluation of the likelihood requires us to regard the hypotheses as considered propositions, not as asserted pro- positions; we can give a definite value to P{'p | q) irrespective of whether q is true or not. This distinction is necessary, because we must be able to consider the consequences of false hypotheses before we can say that they are false. f We get no evidence for a hypothesis by merely working out its consequences and showing that they agree with some observa- tions, because it may happen that a wide range of otlicr hy])ot]ieses would agree with those observations equally well. To get evidence for it we must also examine its various contradictories and show that they do not fit the observations. This elementary principle is often over- looked in alleged scientific work, which proceeds by stating a hyjio- thesis, quoting masses of results of observation that might be expect ed on that hypothesis and possibly on several contradictory ones, ignoring all that would not be expected on it, but might be expected on some alternative, and claiming that the observations support the hypothesis. Most of the current presentations of the theory of relativity (the essen- tials of which are supported by observation) are of this type; so are those of the theory of continental drift (the hypotheses of which are contra- dicted by every other check that has been applied). So long as alter- natives are not examined and compared with the whole of the relevant data, a hypothesis can never be more than a considered one.

In general the probability of an empirical proposition is subject to some considered hypothesis, which usually involves a number of quanti- tative parameters. Besides this, the general principles of the theory and of pure mathematics will be part of the data. It is convenient to have a summary notation for the set of propositions accepted throughout an investigation ; I shall use H to denote it. H will include the specifica- tion of the conditions of an observation. 6 will often be used to denote the observational data.

t This is the reason for rejecting the Principia definition of implication, which leads to the proposition, If is false, then q implies p.’ Thus any observational result p could be regarded as confirming a false hypothesis q. In terms of cntailment the corresponding proposition, ‘If ^ is false, q entails p’, does not hold irrespective of p.

§2.1

DIRECT PROBABILITIES

49

2.1. Sampling. Suppose that we have a population, composed of members of two types <f) and ^ (f>, in known numbers. A sample of given numb('r is drawn in such a way that any set of that number in the population is equally likely to be taken. What, on these data, is the probability that the numbers of the two types will have a given pair of values?

Let r and s be the numbers ol types (/> and in the population, I and rn those in the sample. The number of possible samples, subject to the conditions, is the number of ways of choosing Z+m things from r+s, which we denote by The number of them that will have

precisely I things of type ^ and m of type ^ (j> is <^ata

H any two particular samples are exclusive alternatives and are equally probable; and some sample of total number l-\ ni must occur. Hence the probability that any particular sample will occur is 1 / ; and

the probability that the actual numbers will be I and in is obtained, by the addition rule, by multiplying this by the total number of samples with these numbers. Hence

(1)

It is an easy algebraic exercise to verify tliat the sum of all these ex- pressions for different values of Z, Z+m remaining the same, is 1.

Explicit statement of the data H is desirable because it may be true in some cases that all samples are possible but not equally probable. In such cases the application of the rule may lead to results that are seriously wrong. To obtain a genuine random sample involves indeed a difficult technique. Yule and Kendall give examples of the dangers of supposing that a sample taken without any particular thought is a random sample. They are all rather more complicated than this problem. But the following would illustrate the point. Suppose that we want to know the general opinion of British adults on a political question. The most thorough method would be a referendum to the entire electorate. But a newspaper may attempt to find it by means of a vote among its readers. These will include many regular subscribers, and also many casual purchasers. It is possible that on a given day any individual might obtain the paper even if it was only because all the others were sold out. Thus all the conditions in H are satisfied, except that of randomness; because on the day when the voting-papers are issued there is not an equal chance of a regular subscriber and an occa- sional purchaser obtaining that particular number of the paper. The tendency of such a vote would therefore be to give an excess chance

3505.58 £

60

DIRECT PROBABILITIES

Chap. II

of a sample containing a disproportionately high number of regular subscribers, who would presumably be more in sympathy with the general policy of the paper than the bulk of the population.

2.11. Another type of sampling, which is extensively discussed in the literature, is known as sampling with replacement. In this case every member, after being examined, is replaced before the next draw. At each stage every member, whether previously examined or not, is taken to be equally likely to be drawn at any particular draw. This is not true in simple sampUng, because a member already examined cannot be dravTi at the next draw. If r and s as before are the numbers of the types in the population, the chance at any draw of a member of the first type being drawn, given the results of ail the previous draws, will always be r*/(r+5), and that of one of the second type sl(r-\-s). This problem is a specimen of the cases where the probabilities reduce to chances.

Many other actual cases are chances or approximate to them. Thus the probabilities that a coin will throw a head, or a die a 6, appear to be chances, as far as we can tell at present. This may not be strictly true, however, since either, if thrown a sufficient number of times, would in general wear unevenly, and the probabilit}’' of a head or a six on the next throw, given all previous throws, would depend partly on the amount of this wear, which could be estimated by considering the previous throws. Thus it would not be a chance. The existence of chances in these cases would not assert that the chance of a head is ^ or that of a six the latter indeed seems to be untrue, though it is near enough for most practical purposes.

If the chance of an event of the first type (which we may now call a success) is x, and that of one of the second, which we shall call a failure, is l—x y, then the joint probability that i+m trials will give just I successes and m failures, in any prescribed order, is x^y^. But there will be ^+^1 ways of assigning the I successes to possible positions in the series, and these are all equally probable. Hence in this case

(2)

V ! Tit I

which is a typical term in the binomial expression for (x+yy+”^. Hence this law is usually known as the binomial distribution. In the case of sampling with replacement it becomes

P{l,m \H) ~

(l+m)\l r VI s l\m\ \r+5/\r+5/

(3)

§2.1

DIRECT PROBABILITIES

51

It 18 easy to verify that with either type of sampling the most probable value of I is within one unit of r(l+m)j(r-\-8), so that the ratio of the types in the sample is approximately the ratio in the population sampled. This may be expressed by saying that in the conditions of random sampling or sampling with replacement the most probable sample is a fair one. It can also be shown easily that if we consider in succession larger and larger populations sampled, the size of the sample always remaining the same, but r and s tending to infinity in such a way that rjs tends to a fixed value xjy, the formula for simple sampling tends to the binomial one. What this means is that if the population is sufficiently large compared with the sample, the extraction of the sample makes a negligible difference to the probability at the next trial, which can therefore be regarded as a chance with sufficient accuracy.

2.12. Consider now what happens to the binomial law when I and m are large and x fixed. Let us put

/(!) - i!»!

}

(4)

l-\-m ~ /i; 1 nx-~\-n'~^^oi;

m ny—n^^^oLf

(5)

and suppose that a. is not large. Then

^ogfil) ~ logZ!+logm!— logn!

llogx—mlogy.

(6)

Now we have Stirling’s formulaf

logn! = (TO+pogn-w+41og27r+^-o|lj.

(7)

Substituting and neglecting terms of order 1/i, 1/m, we have

log/(;) - ilog^^+Zlog^ + wlog^.

n nx ny

(8)

t The closeness of Stirling's approximation, even if l/12n is neglected, is remarkable. Thus for n = 1 and 2 it gives

1! = 0-9221; 2! = 1-9190;

while if the term in l/12n is kept it gives

1! = 1-0022; 2! =3 2-0006.

Considered as approximations on the hypothesis that 1 and 2 are large numbers they are very creditable. The use of the logarithmic series may lead to larger errors.

Proofs of the formula and of other properties of the factorial function, not restricted to integral argument, are given in H, and B. S. Jeffreys, Methods of Mathematical Physics, Chapter 15.

DIRECT PROBABILITIES

62

Chap. II

Now substituting for I and m, and expanding the logarithms to order a* we have

log/(0 = \^og{27rnxy) + -^ + 0(oLH-^f\ (9)

1.1 ( (l-^nxf]

f(l) * {27rnxyyi‘^^^^\ 2nxy )

This form is due to De Moivre.f From inspection of the terras neglected we see that this will be a good approximation if I and m are large and a not large compared with or Also if nxy is large the chance varies little between consecutive values of Z, and the sura over a range of values may be closely replaced by an integral, which will be valid as an approximation till I— nx is more than a few times (nxyY^'K But the integrand falls off with l—nx so rapidly that the integral over the range where (10) is valid is practically 1, and therefore includes nearly all the chance. But the whole probability of all values of Hs 1 . It follows that nearly the whole probability of values of I is concentrated in a range such that (10) is a good approximation to (4).

It follows further that if we choose any two positive numbers ^ and y, and consider the probability that I will lie between ??(a:+^) and n{x—y), it will be approximately

-y

which, if p and y remain fixed, will tend to 1 as tends to infinity. That is, the probability that (l—nx)!n will lie within any specified limits, however close, provided that they are of opposite signs, will tend to certainty.

2.13. This theorem was given by James Bernoulli in the Ars Conje- ctandi (1713). It is sometimes known as the law of averages or the law of large numbers. It is an important theorem, though it has often been misinterpreted. We must notice that it does not prove that the ratio l/n mil tend to limit x when n tends to infinity. It proves that, subject to the probability at every trial remaining the same, however many trials we make, and whatever the results of previous trials, we may reasonably expect that Ijn—x will lie within any specified range about 0 for any particular value of n greater than some assignable one depending on this range. The larger n is, the more closely will this probability approach to certainty, tending to 1 in the limit . The

t Miscellanea Analytical 1733.

§2.1

DIRECT PROBABILITIES

53

existence of a limit for Ijn would require that there shall be a series of positive numbers depending on n and tending to 0 as n -> sucli that, for all values of n greater than some specified Hq, Ijn—x lies between ^^Rt it cannot be proved mathematically that such series

always exist when tlie sampling is random. Indeed we can produce possible results of random sampling where they do not exist. Suppose that X - . it is essential to the notion of randomness that the results

of previous trials are iiTelovant to the next. Consequently we can never say at any definite stage that a particular result is out of the question. Thus if we enter 1 for each success and 0 for each failure such series as the following could arise:

1001 100101001001 1 1010..., 10010010010010010 0100...,

0 000000 0 0 000000000000...,

111111111111111111111...,

101 100001 1 I 1 1 1 1 10000000000....

Tlie first series was obtained by tossing a coin. The others were systematically designed; but it is impossible to say logically at any stage that t lie conditions of the problem forbid the alternative chosen. They are all possible results of random sampling consistent with a chance i. But the second would give limit the third and fourth limits 0 and 1 ; tlie fifth would give no limit at all, the ratio Ijn oscil- lating between J and |. (The rule adopted for this is that the number of zeros or units in each block is equal to the whole number of figures before tlie beginning of the block.) An infinite number of series could be chosen that would all be possible results of random selection, assuming an infinite number of random selections possible at all, and giving either a limit different from i or no limit.

It was proved by Wrinch and me,’j* and another version of the proof is given by M. S. Bartlett, | that if we take di fixed a independent of 7^,, Uq can always be cliosen so that the probability that there will be no deviation numerically greater than a, for any n greater than is as near 1 as we like. But since the required tends to infinity as <x tends to 0, we have the phenomenon of convergence with infinite slowness that led to the introduction of the notion of uniform convergence. It is necessary, to prove the convergence of the series, that shall tend to 0; it must not be independent of n, otherwise the ratio might oscillate finitely for ever.

t Phil. Mag. 38, 1919, 718-19. J Proc. Boy. Boc. A, 141, 1933, 620-1.

54

DIRECT PROBABILITIES

Chap. II

Before considering this further we need a pair of bounds for the incomplete factorial function,

CO

(1)

(2)

I = j du,

X

where x is large. Then

00

/ > X” J e-*^du =

X

Also, if w = x-\-Vj

ujx < expv/x,

1 < x^e~^^ I expi —t-^~}vdv = .

j \ x/ t—nx

Hence, if xjn is large,

(3)

(4)

Now let P{n) be the chance of a ratio in n trials outside the range x±:oc. This is asymptotically

by putting oc^ = u and applying (4).

Now take (6)

The total chance that there will be a deviation greater than for some n greater than is less than the sum of the chances for the separate n, since the alternatives are not exclusive. Hence this chance

W = 7lo

(7)

Put

then

n =

e(n.) <

Vno

^ 2{2x(l-x)}’'» r ni'’ 1

(8)

§2.1

DIRECT PROBABILITIES

65

with a correcting term small compared with the first for large Hence Q{n^ does tend to zero as Uq tends to infinity, and we have the result that Uq can be fixed so that the total chance of deviations greater than for all n greater than is as small as we please; and if all deviations are less than the series converges. Hence it may be expected, with an arbitrarily close approach to certainty, that subject to the conditions of random sampling the ratio in the series will tend to x as a limit. f

This, however, is still a probability theorem and not a mathematically proved one; the mathematical theorem, that the limit must exist in any case, is false because exceptions that are possible in the conditions of random sampling can be stated.

The situation is that the proposition that the ratio does not tend to limit X has probability 0 in the conditions stated. This, however, does not entail that it will tend to this limit. We have seen (1) that series such that the ratio does not tend to limit x are possible in the conditions of the problem, (2) that though a proposition impossible on the data must have probability 0 on those data, the converse is not true; a proposition can have probability 0 and yet be possible in much simpler cases than this, if we maintain Axiom 5, that probabilities on given data form a set of not higher ordinal type than the continuum. If a magnitude, hmited to a continuous set of positive values, is less than any assignable positive quantity, then it is 0, But this is not a contra- diction because the converse of Theorem 2 is false. We need only distinguish between propositions logically contradicted by the data, in which case the impossibility can be proved by the methods of deduc- tive logic, and propositions possible on the data but whose probability is zero, such as that a quantity with a uniform distribution of its prob- ability between 0 and 1 is exactly

The result is not of much practical importance; we never have to count an infinite series empirically given, and though we might like to make inferences about such series we must remember the condition required by Bernoulli’s theorem, that no number of trials, however large, can possibly tell us anything about their immediate successor that we did not know at the outset. It seems that in physical conditions something analogous to the wear of a coin would always violate this condition. Consequently it appears that the problem could never arise. Further, there is a logical difficulty about whether the limit of a ratio

t Another proof is given by F. P. CantelU, Rend. d. circ. maUm.^ Palermo, 41, 1916, 19i~201 ; Rend. d. R. Acad. d. Lincei, 26, 1917, 39-45. See E. C, Fieller, J. R. Stat. Soc. 99, 1936. 717.

56

DIRECT PROBABILITIES

Chap. II

in a random series has any meaning at all. In the infinite series con- sidered in mathematics a law connecting the terms is alwa^^s given, and the sum of any number of terms can be calculated by simply following rules stated at the start. If no such law is given, which is the essence of a random process, there is no means of calculation. The difficulty is associated with what is called the Multiplicative Axiom; this asserts that such a rule always exists, but it has not been proved from the other axioms of mathematical logic, though it has recently been proved by Godel to be consistent with them. Littlewoodf remarks, 'Reflection makes the intuition of its truth doubtful, analysing it into prejudices derived from the finite case, and short of intuition there seems to be nothing in its favour.’ The physical difliculty may arise in a finite number of trials, so that there is no objection to sup- posing that it may arise in any case even if the Multiplicative Axiom is true. In fact I should say that the notion of chance is never more than a considered hypothesis that we are at full liberty to reject. Its useful- ness is not that chances ever exist, but that it is sufficiently precisely stated to lead to inferences definite enough to be tested, and when it is found wrong we shall in the process find out how much it is wrong.

2.14. We can use the actual formula 2.12 (10) to obtain an approxi- mation to the formula for simple sampling when /, m, r~~l, and .s—m are all large. Consider the expression

F -= X , ( 1 )

where x and y are two arbitrary numbers subject to x-\-y -- 1. r, s, and l~\-7n are fixed. Choose x so that the maxima of the two expres- sions multiplied are at the same value of and call this value 1^ and the corresponding value of m, ttIq. Then

Iq r=z rx\ t—Iq = ry; sx\ sy\ (2)

whence (r-f 5)a; = ^o+^o (3)

Then, by 2.12 (10),

F 4= (27Trxy)~’/-exp

(27rxy)-'^{rsy^l‘^exm

\ 2szy I

Also

2rszy f

0 = = {2Tr{r+s)xy}--yK

(4)

(6)

f Elements of the Theory of Real Functions^ 1926, p. 25.

§2.1

DIRECT PROBABILITIES

67

Hence by division

rn sn

"iTrrsxyj

exp -

{l-l,Y{r+s)

whence

P{hm\H)

where

(r-\-sYxy ~ (7)

(,+,.)3 yk I \

‘2Trrs(l-{-7n)(r+s—l—m)j " [ 2rs{l-\-m)(r-j-s l—tn.)j’

(8)

_r{l+m)

Comparing this with 2.12 (10) we see that it is of similar form, and the same considerations about the treatment of the tail will apply. If r and s are very large compared with I and m, we can write

r ^ (r+s)p; s ^ (r+s)q, (10)

p and q now corresponding to the x and y of the binomial law^ ; and the result approximates to

I !

\27T(l+m)pql 2(l-i-m)pq J'

which is equivalent to 2. 12(10). Tn this form we see that the probabilities of different compositions of the sample depend only on the sample and on the ratio of the type numbers in the population sampled; provided that the population is large compared with the sample, further informa- tion about its size is practically irrelevant. But in general, on account of the factor (r+t9)/(r+5— m) in the exponent, the probability will be somewhat more closely concentrated about the maximum than for the corresponding binomial. This represents the effect of the with- drawal of the first parts of the sample on the probabilities of the later parts, which will have a tendency to correct any departure from fairness in the earlier ones.

2.15. Multiple sampling and the multinomial law. These are straightforward extensions of the laws for simple sampling and the binomial law. In the first case, the population consists of p different types instead of two, the numbers being r^, rg,.-, the corresponding numbers in the sample are 7ij, ng,..., n^j with a prescribed total. It is supposed as before that all possible samples of the given total number are equally probable. The result is

(1)

58

DIRECT PROBABILITIES

Chap. II

In the second case, the chances of the respective types occurring at any trial are (their total being 1) and the number of trials

2 w is prescribed. The result is

^1* ^2* ^jo- lt is easy to verify in ( 1 ) that the most probable set of values of the n *s are nearly in the ratios of the r’s, and in (2) that the most probable set are nearly in the ratios of the x's. Consequently we may in both cases speak of the expected or calculated values; if JV is the prescribed total number of the sample, the expected for multiple sampling will be NrJ'^ r, and the expected for the multinomial will be Nx^. The probability will, however, in both cases be spread over a range about the most probable values, and we shall need to attend later to the question of how great a departure from the most probable values, on the hypo- thesis we are considering, can be tolerated before we can say that there is evidence against the hypothesis.

2.16. The Poisson law.f We have seen that the use of Stirling’s formula in the approximation used for the binomial law involves the neglect of terms of order 1/Z and 1/m, while the result shows that there is a considerable probability of departures of I from nx of amounts of order {nxyY^^. If then (nxyf^ > nx, the result shows that i = 0 is a very probable value, and the approximation must fail. But if n is large, this condition implies that x is small enough for nx to be less than 1 . Special attention is therefore needed to cases where n is large but nx moderate. We take the binomial law in the form

Now log{»!/(w— Z)!} = nogn4-0(Z*/ra). (2)

Also, since x is small, (1— x)™-* = (3)

nearly; whence, so long as l^jn and lx are small,

= (4)

The sum of this for all values of Z is unity, the terms being times the terms of the expansion of c“*. The formula is the limit of the binomial when n tends to infinity and a: to 0, but tix to a definite value. If nx* is small but nx large, both approximations to the binomial are valid.

The condition for the Poisson law is that there shall be a small chance

t S. D. Poisson, Recherches sur la probabilite des jugements, 1837, pp. 206-7.

§2.1

DIRECT PROBABILITIES

69

of an event in any one trial, but there are so many trials that there is an appreciable probability that the event will occur in some of them. One of the best-known cases is the study of von Bortkiewicz on the number of men killed by the kick of a horse in certain Prussian army corps in twenty years. The unit being one army corps for one year, the data for fourteen corps for twenty years gave the following summary. f

Number of deaths

Number of units

Expected

0

144

1390

1

91

97-3

2

32

34- 1

3

11

80

4

2

1-4

5 and more

0

0*2

The analysis here would be that the chance of any one man being killed by a horse in a year is small, but the number of men in an army corps is such that the chance that there will be one man killed in an entire corps is appreciable. The probabilities that there will be 0, 1, 2,... men killed in a corps in a year are therefore given by the Poisson rule; and then by the multinomial rule, in a sample of 280 units, we should expect the observed numbers to be in approximately the ratios of these probabilities. The column headed ‘expected’ gives the expectations on the hypothesis that nx 0*70. They have been recalculated, the calculated values as quoted having been derived from several Poisson laws superposed.

Another instance is radioactive disintegration. The chance of a particular atom of a radioactive element breaking up in a given interval may he very small; but a specimen of the substance may contain something of the order of 10^® atoms, and the chance that some of them may break up is appreciable. The following table, due to Rutherford and Geiger, J gives the observed and expected numbers of intervals of J minute when 0, 1, 2,... a-particles were ejected by a specimen.

Number 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14

Obs. 67 203 383 525 632 408 273 139 46 27 10 4 0 1 1

Exp. 54 211 407 626 508 393 254 140 68 29 11 4 1 0 0

0-E +3 -8 -24 0 +24 +15 +19 -1 ~23 -2 -1 0 -1 +1 +1

nx is taken as the total number of particles divided by the total number of intervals = 10097/2608 == 3*87. It is clear that the Poisson law agrees with the observed variation within about one-twentieth of its range; a closer check will be given later.

t von Bortkiewicz, Das Gesetz d. kleinen Zahlerif 1898. Quoted by Keynes, p. 402.

X Rutherford. H. Geiger, and H. Bateman, Phil. Mag. 20, 1910, 698-707.

60

DIRECT PROBABILITIES

Chap. II

The Aitken dust-counter provides an example from meteorology. f The problem is to estimate the number of dust nuclei in the air. A known volume of air is admitted into a chamber containing moisture and filtered air, and is then made to expand. This causes condensation to take place on the nuclei. The drops in a small volume fall on to a stage and are counted. Here the large number is the number of nuclei in the chamber, the small chance is the chance that any particular one will be within the small volume at the moment of sampling. Scrase gives the following values.

Number

0

1

2

. J

4

5

G

7

S

Ohs,

23

56

88

95

73

40

17

5

3

Exp.

25

65

88

82

61

38

21

10

4

0-E

2

~9

0

f 13

+ 12

+ 2

-4

5

-1

The data are not homogeneous, the observations having been made on twenty different days; nxwebs estimated separately for each and the separ- ate expectations were calculated and added. It appears that the method gives a fair representation of the observed counts, though there are signs of a systematic departure. Scrase suggests that in some cases zero counts may have been wrongly rejected under the impression that the instrument was not working. This would lead to an overestimate of 7ix on some days, therefore to an overestimate of the expectations for large numbers, and therefore to negative residuals at the right of the table. Mr. Diananda points out that the observed counts agree quite well with nx - 2*925.

2.2. The normal law of error. Let us suppose that a quantity that we are trying to measure is equal to A, but that there are various pos- sible disturbances, n in number, each of wliicli in any particular case has equal chances | of producing alterations in the actual measure; the sign of the contribution from each is independent of those of the others. This is a case of the binomial law. If I of the components in an individual observation are positive and the remaining n—l negative, the measured value will be

X A-f-/e— (/i— /)€ = A-f-(2Z— 7^)e. (1)

The possible measured values will then differ from A— Tie by even multiples of e. We suppose n large. Then the probabilities of different values of I are distributed according to the law obtained by putting X y r- I in 2.12 (10), namely,

(2)

t John Aitkon, Proc. Roy, Soc. Edin, 16, 1888, 135-72; F. J. Scraae, Q.J.R, Met, Soc, 61, 1935, 368-78.

§2.2

DIRECT PROBABILITIES

61

and the probability that I will be equal to Zj, (> Zj), or some inter- mediate value will be

= I (3)

l=^ll

But this is the probability that the measure x will be in the range from A-f (2Zj--7i)€ to A+(2Z2— n)e, inclusive. If, then, we consider a range to :r2, long enough to include many possible values of Z, we can replace the sum by an integral, write

I— In = (a:-~A)/2€, (4)

and (5)

This range will contain (.^2— a;i)/2e-f 1 admissible values of x. Now

suppose that Xg— which is much larger than e, is also much less than cVn. The sum will then approximate to

(6)

Now let n be very large and e very small, in such a way that is finite. The possible values of x will then become indefinitely closely packed, and if we now consider a small range from Xj^ to the chance that

X lies within it will approximate to

P(x^ < X < So: I H)

1

{27rnf‘^€

2n€^ ]

hx.

(7)

This is an instance of the normal law, which we can WTite in its general

1 I It

P(x.<x<x.+dx|fl)_^expj-LL^jfe ,8) or, more briefly,

in the sense that w^hen dx tends to zero the ratio of the two sides tends to 1. In practice we are always concerned with finite ranges, so that strictly we alw ays require the integrals of these expressions over some finite range, and the transition from Sx to dx involves only a step that we shall always undo before we make any use of th^ results.

It will be noticed that whereas we started with three parameters A, n, and €, in the result we are left with two, A and cVti, the latter being replaced by cr. This is similar to what happens in sampling, where the

62

DIRECT PROBABILITIES

Chap. II

size of the population sampled becomes irrelevant when it is large. The form of the normal law, in application to errors, seems to have been given first by Laplace in 1783, though it is usually attributed to Gauss. f The law can also be written

P{x-^ < X < x^+dx \H) = ^exp{—h^{x—X)^}dx, (9)

where 2AV 1. (10)

a is usually called the standard error, but sometimes the mean square error or simply the mean error, h is called the precision constant. If we introduce the error function

X

evix f e-^dt. (11)

VTT J 0

the probability that x will be less than x^ is ^{14-erfA(Xi— A)}. Tables of the probability that x—X will be less than given multiples of a are given by Sheppard and by later writers. The error function, which has other applications in heat conduction and diffusion, is tabulated by Milne-Thomson and Comrie. In statistical applications (8) is more convenient than (11), since a usually arises more directly than h. The curve j/oc exp{—(x—X)^l2a^} has inflexions at A±a. There is a prob- ability 0*683 that an observation will lie between A±(t. There is a probability | that it will lie between A±0*6745cr. In this sense 0*6745(7 is often called the probable error, and is the uncertainty usually quoted in astronomical and physical works. This practice would be better abandoned. In applying any significance test or the ^ or t rules what arises is a, and if uncertainties are given in terms of the probable error, the multiplication must first be undone, with unnecessary trouble and some loss of accuracy due to accumulation of rounding-off errors.

The conditions contemplated in the normal law of error have often a rough justification. In many cases we have adequate reason to suppose that the quantity we are trying to measure has a ‘true value’, though we must reserve a further discussion of what that can mean in relation to our general theory. But several minor disturbances may affect any individual measure, such as wandering of the observer’s attention, the fact that he must round off his measures to the nearest multiple or tenth of the scale interval, disturbance of the apparatus through vibra- tion of the ground or wind, and so on. These can often be regarded as independent. They are not in general capable of producing only two t Pearson, Biometrika, 13, 1920, 25.

§2.2

DIRECT PROBABILITIES

63

equal and opposite values of the disturbance; most of them are capable of a continuous range of values, and in general there is not much reason to suppose that these are equally spread for all the disturbances. The general application of the above argument must therefore be mistrusted. It can be regarded only as an indication that there may be cases where the chance of error is distributed according to the normal law, which sums up the whole information with regard to the possible variation in two parameters A and a, A is also often caUed the population mean and a the population standard deviation. The latter term is rather cumbrous, and if the word ‘population’ is omitted it is liable to be confused with the standard deviation of a given finite set of observations, which is not the same thing.

Where we are dealing with a law of the form

of which the normal law is an instance, we may speak of A as the location parameter and o as the scale parameter, to use Fisher’s terms. These correspond to epistemological needs better than ‘true value’ and ‘standard error’ do. But the latter terms are convenient; we have only to remember that ‘true value’ is not to be understood in an absolute sense, but in the sense that any law relating measures, if it is to be of any use, must be clearly stated, in probability terms, and that a possible way of progress (apparently the only possible way) is to treat the variation as the resultant of a part that would be exactly predictable, given exact statements of the values of certain parameters, and a random error. The law in its naive form w^ould deal only with the former part. The parameters in this part may be called the true values of the parameters, and the observed values that they would lead to if the random part was neglected the true values. The actual observed values will differ somewhat. By the principle of inverse probability we shall be able then to proceed from the observations to estimates of the true values of the parameters, which, however, will not be exact deter* minations, but will have ranges of uncertainty corresponding to the fact that the individual random errors in the observations are not definitely known.

In actual fact there are some cases where the normal law of error appears to represent the outstanding variation as well as we can tell. There are others where, though we find that it is probably incorrect when we study a sufficient number of observations, this number is

64

DIRECT PROBABILITIES

Chap. II

large, of the order of 500, and the use of the normal law in such cases as if it was correct would not lead to serious mistakes. There are others where it is glaringly wrong, and the only proper treatment is to obtain a sufficient number of observations to give us some idea of what the corresponding distribution of chance can be. Meanwhile we shall con- sider an important series of generalized laws of error.

2.3. The Pearson laws. If we write the normal law of error in the form

where we liave now made the parameters A and a explicit (they were formerly understood in H), we see that it is an instance of the general

P(dx\H) = ydx, (2)

where y ^ 0 and the integral of y over all possible values must be 1 . In this case we find easily

(3)

y ax

The law, therefore, has the properties that dyjdx vanishes in the limit when y tends to 0, and at one intermediate value of x, namely, A. If we consider the generalized form

(4)

y dx bQ-\~b^x~\-b2X^'

the same will usually hold, but e have two more parameters and shall be able to represent laws of a much wider range of form. They will have one point where y is stationary; if the range of x is infinite y and dyjdx will tend to zero at the end or ends; if the range is limited in one or both directions there will still be cases where this holds. The integral of (4) can in general be written in the form

y ^ A{x~c^)^^(c^-xf\ (5)

where A will be fixed by the condition that the integral of i/ is 1 , and Cj and Cg are the zeros of the denominator in (4). There are three main types of solution and a number of transitional and degenerate cases.

1. q and Cg imaginary. Then they must be conjugate complexes, and for y to be real and mg must also be conjugate complexes. y cannot vanish or become infinite for any real value of x, and the admissible values of x range from oo to +oo, with a maximum of y

§2.3

DIRECT PROBABILITIES

65

at some intermediate value. Forms with one maximum are designated bell-shaped by Pearson. We may wTitc these laws in the forms

m

2m-i (m—l-l-ig)! (in- ]-~iq )[

27T(2m— 2)1

X exp|— 2gtan“i^^j. (6)

These are Pearson’s Type IV. In general they are asymmetrical or skew, but if ^ = 0 they reduce to the symmetrical form

T(2m-~2)\

'2W/-1

{7n— 1 )! 7T^/'^{ni~~ I ) !

(7)

(8)

which is Pearson’s Type V 1 1. In both cases m must be greater than | for convergence. These law’s resemble tlie normal law- in having an infinite range of x in both directions, w hich is true of no other Pearson type, but y falls off less rapidly. With the normal law" the expectation of any power of x is finite; with Type VII that of any even power equal to 2m— 1 or more is infinite (m- need not be integral); w ith Type IV expectations of odd pow ers V 2m 1 are also infinite. This is a useful property in representing errors of measurement, since it is usually found, w"hen sufficient observations are available, that there are more outlying large residuals than the normal law" w ould suggest. The fact that these law s, like the normal law’, give a non -zero chance of an error greater than any finite amount is an apjjarent draw^back, since we might say that however bad the observations are there is some hmit to the error; l)ut to harmonize this belief wdth the observed distributions would require us to go beyond the range of the Pearson types, which do give satis- factory agreement withm the ranges where observations exist.

If Cj and Cg are real (Cg > c^) we must distinguisli three cases. (4) has singularities at and Cg and the solution is applicable only in ranges that do not include a singularity. Hence we must consider separately cases where the admissible values of x are less than c^, between and Cg, or greater than Cg. The difference between the first and third can be removed by merely reversing the direction of measurement.

2. Admissible values of x betw een and Cg. We can take the law in

the form

y ^

(m^4-mg+l)!

m^! mg! (Cg— Cj)"'! ^"'3+1

(9)

F

3595.58

DIRECT PROBABILITIES

Chap. II

which will be possible if both and r/ig are greater than —1. If both are positive, the curve is belhshaped. If 0 > > 1, 2/ infinite at

Cj. If at the same time is positive, dyjdx is negative throughout the range and the curve is called d -shaped. In this case a does not lie between Cj and and is not an admissible value of x. If and are both negative, y is infinite at both limits and a lies between them. The curve is then called \J -shaped. These cases cover Pearson’s Type I, It will be seen that the possibility of U-shaped and J -shaped curves gives it greater generality than was originally attempted.

There are several special cases:

= Wig. The law is then symmetrical. This is Pearson’s Type II.

Further degenerations give

7^2 ~ m2 0. This makes y uniform between and Cg, and zero outside that range. This is the rectangular distribution, not given a number by Pearson.

= mg 1. This, with a change of scale and origin, gives yoc 1 ~x^, the parabolic distribution,

mj ~ 0. This is a J -shaped curve with y proportional to {c^—x)^^ for X between and c^. This is Pearson’s Type IX. It starts from a finite ordinate at Cj.

1 j with 1 < m < 1 .

This is Pearson's Type XI 1, The curve is always J -shaped.

3. Admissible values of x all 5^ Cg. We can take the law in the form

y =

(-mj— 1)!

m^\{-

-rn^

-mg— 2)! (Cg—

(x-c^)m(C2~-x)^\ (10)

where for convergence m,g > —1, < ~1. These are the laws

of Type VI. If mg 0 they are bell-shaped, if m^g 0, J -shaped. They are never U-shaped. These laws will give the kind of distribution shown by the times of arrival of a train; there is a concentration at values a little greater than Cg, values less than Cg do not occur, and there is a long train of large values, which may rarely occur but are serious when they do.

A particular case is

mg = 0. This makes y proportional to for values of x greater

than Cg; evidently m2 < 1. This gives Pearson’s Types VIII and XI, which are identical. It starts from a finite ordinate at Cg.

§2.3

DIRECT PROBABILITIES

67

lYpes IV, I, and VI, to take them in what seems to me to be their natural order, are the only ones that involve the full number of adjust- able parameters, four. There are also three transitional cases between them.

4. There will be a transition from Type I to Type VI expressed by

making Cg in I tend to +cc or in VI to oo. In either case the hmiting

form IS a > 0).

This is Type 111. It resembles Type VI in appearance but is more closely concentrated to small departures from c. A particular case is

m ~ 0; this is Type X, an exponential la^v, which can also be regarded as the transition between Types VIII and IX.

5. The transition from Type VI to Type IV is the case of equal roots, the roots of the denominator in (4) being equal, real, and finite. Then we can write (4) in the form

\ dy ^ ^

y dx x—c'(x~cy^'

whence y = A(.r--c)“''exp

This is Type V. To give convergence at x, a must be > 1; for con- vergence at c, ^ > 0 for any x > It is always bell-shaped, since y must vanish at x c. Othei^vise it resembles Type VI. It differs from Type III in the interchange of the two types of convergence at the extremes; indeed, the change of (x—c) to {x~c)~^ transforms one into the other.

6. The transition from Type IV to Type I requires the roots to be ±x ; then and both vanish and we are back to the normal law.

This analysis covers the range of the Pearson types, and is, I think, considerably shorter and more systematic than has been given pre- viously. jNIy own experience with them has been rather small, though I have had to deal with Types II, III, VII, and VIII. For purposes of exposition I think it would be a great convenience if those who use them extensively could agree on a more systematic numbering in place of the present haphazard one, which places III, the transition between I and VI, between II, which is the symmetrical case of I, and IV, which is a different main type from any; and VI, a main type, between V, a transitional case, and VII, a degenerate case of IV. I should suggest the following.

68

DIRECT

PROBABILITIES

Chap. II

A umber

Main types

Pearson's number

Special rases

Pearson 's

Suggested

1

IV

- 0

vn

\a

2

J

7??1

- m2

n

2a

7/1,

- //?2 0

Roct.

2b

'//?,

//?2 I

Parab.

2c

»?1

0 "

I.\

2d

'//I,

//?,,

Xll

2e

3

VI

m2

-- 0

VJli

:\a

Transitions

2 to 3

JIl

m

0

X

2\\a

3 to I

V

1 to 2

Xonnal

This covers the whole range witli the exception of XI, which is a mere rewriting of VIll. I think tliat special numbers for the rectangular ami parabolic laws are worth while as they are likely to be at k'ast as im- portant as Xll in practice, and the rt^ctangular law has great theoretical interest. Both, like t he normal law. involve only a scale parameter and a location parameter. Tlie main types involve iwo others. The rest involve three parameters in all.

It may be remarked that JVarson distinguislual Types 1 and VI according as the roots are real and of ()p])osite sign or real and of like sign. This appears to make the type depend on the arbitrary position of the origin. The important point is whether the admissible values of X lie between the roots or not. In fact I^earson does make his decision according to the latter criterion.

2.4. The negative binomial law. 8u])pose that a distribution of chance follows the Poisson law

P(l\rU) = ^e-r (1)

but that r itself is unknown, having a distribution of chance given by the Type III law

P{dr\H) = (2)

a!

(where, since a may be fractional, we must understand a! to be defined

00

by a! = J Then

0

P{1, dr I H) - ^3y’

To get the total probability for any value of I, we must add for all

§2.4

DIRECT PROBABILITIES

69

possible values o( r\ wliicli means in this case that we must integrate. Then

mu)^- ,4)

J Hn! (]J /!)'»•■'/!»! ' '

0

Apart from the factor (-J~\

^ 1^+1/

expansion of

this is the coefficient of in the

sum over all values of / is 1, as it must be since the conditions stated are exhaustive. If we ])ut

^ + 1

1— a.

we have

(5)

which puts the negative binomial form more cieai ly in evidence. This result is due to M. Greenwood and (h U. Yule. I The immediate a])])lica- tion was to problems of factory accidents. The conditions of tlie Poisson law were satisfied in res[)ect ol‘ the total chance of an accident in a factory in a given period being the sum of a large number of small chances, but it was not clear that these chances were the same for all employees. The chance of a particular workman liaving an accident on a particular day, for instance, would have to be regarded as the analogue of X in the derivation of the Poisson law, and tlic number of days in the period considered as the analogue of n. Then for each individual the chances of 0, 1, 2,... accidents in the period would follow a Poisson law subject to the condition that having one accident does not stimulate him to have another and if the values of r ^ 7ix for the different v ork- men are distributed, as nearlj^ as can be for a finite number, in a Type III law, the negative binomial follows as the resultant for all wnrkmen.

The following alternative development shows that the condition that the probabihties of accidents to the same workman must be independent is not strictly necessary. It can at any rate be replaced by other condi- tions. Suppose that the total number of events is recorded, but that in fact some of the events are composite, two or more being associated. These are each only one independent event, but will be counted as two or more each in the totals. Let r-g,... be the appropriate values of r for the simple, double,... events in the interval considered. Each type

t J. li. Slat. Soc, 83, 1920, 255-79.

70

DIRECT PROBABILITIES Chap. II

separately will satisfy the Poisson rule, and the chance that there will be simple, in^ double events, and so on, will be

|r,,r2 R) = -i--; -^...exp{-(r,+r2+...)}. (6)

The probability that the total number of events as counted will be m is the sum of these expressions, subject to

+ ... 7n. (7)

But this sum is the coefficient of in the expansion of

f{x) - exp(riX+r2x24-...-ri-r2~...). (8)

Now in practice, if we have no record of the individual events, there will not be much hope of determining the separately. But if we want to find a law that will take into account the extra complication we must have at least one new parameter, though there may not be much point in introducing more than one. Let us take the form:

r^--= r^a^-^js. (9)

\ogf{x) = r,^(l + |aar+iaV-i-.-.)— ia+...)

('•i/«){-log(l-aj;) + log(l-a.)}, f{x)

\l~axl

and the coefficient of is

P(m I r^.a,H) l]^,

a\a I \(i /ml

(10)

(H)

(12)

which again is a negative binomial law, with r^ja replacing the a+1 of Greenwood and Yule’s derivation. f

It is convenient to take the law in the form

P(m I r,n,H)

' n \”7i(n+l)...(n-f 1)/ ^ n-\~rj ml \n+r/

(13)

When n -> 00 this tends to the Poisson law with parameter r. We shall see later that it has other advantages. The series converges for all positive n. The expectations of m and are r and {l~\-\jn)r^.

That of (m~~r)^ is r-\-r^/7t. When n-> 0, all the chances of non-zero m tend to 0, while that of m being zero tends to 1. In the latter case as we approach the limit, keeping r fixed, the chances of m become more and more widely spread to wide values, and the concentration at 0 is needed to keep the total expectation equal to r. Thus the negative

t This derivation has already been given by R. Liiders, Biometrika^ 26, 1934, 108-28.

DIRECT PROBABILITIES

71

§ 2A

binomial law, for small n, will resemble the distribution of the scores of a first-class cricket or billiards player, whose commonest score may be 0 though his average is about 60. On the Poisson law the commonest score and the average should approximately agree, and the cliance of a score of 1 would be 60 times that of a score 0.

Here we have a case where two different types of departure from tlie Poisson law both lead to results of the same form, and modify it in the same direction. If the law is nevertheless found to agree with the facts, it is reasonable to reject both types of departure. Thus the agreement of the data about deaths from kicks of a horse in the Prussian army may be taken to mean both (1) that nobody can be killed twice by the kick of a horse, (2) that the fact that one man has been so killed does not indicate an extra liability for others in the same unit to be. The agreement in the radioactivity data would mean that (1) the chances of disintegration of different atoms of the same radioactive substance are approximately equal, (2) the disintegration of one atom does not lead immediately to the disintegration of another.

2.5. Correlation. This can be treated on lines analogous to the deriva- tion of the normal law from the binomial. Suppose that two quantities X and y are to be measured simultaneously, and that there are m-\-n independent component variations, each contributing to a:: and to y. m of them are constrained to give the same sign in both x and y, n to give opposite signs. Suppose that in a particular case the number making positive contributions to x that give the same sign is p, the number giving opposite signs q. Then

X pa— (m—pja+ya— (n— = (2p~7n)(x-\-(2q—n)cx, (1)

y = pP-{ni—p)P~qp+{n-q)p = (2p—m)P-{2q—n)p. (2)

We are taking each component to be as likely as not to give a positive contribution to x. Then

P(p, q\m,n, ot, p, H) = HJj, ”(7,

by the previous argument. We have to transform to the observed variables x and y. Now

8{p,q) ^ .2\a /?/

(4)

Remembering that p and q are capable of integral values only, and that

72

DIRECT PROBABILITIES

Chap. II

the total chance in any region must be the same whether the observation is expressed in terms of p and g or of x and y, we see that we must replace the sum with regard to p and q by the integral with regard to dxdyj^cxp. Hence

P(dxdy I in, n, a, H)

dxdy ^ f 1 /^ , ^

^ I Sm\a 8?l\ot ,

(5)

Now put

{7n+n)or a-; (m+n)/3^ t“;

(m n)cxP ~ par. (6)

Then we find P{dxdy I m, n, a. jS, II)

dxdy

{ 1 Iz^

exp -^71 -i

I ^(1— P )\«^

(7)

SO that the four original })arametcrs are now reduced to three, and we can assert that this is also equal to P{dxdy | a, r.p. H). Of course, every- thing that can b(‘ said against the normal law of error for one variable can be said twice against this form, which is the generalization to two variables. But on the other hand the chief thing that can be said in favour of the normal lawq that of all laws that are anywhere near the truth it is far tlie easiest to apply, can also be said wdth greater force of normal correlation. The ne^v parameter p is called the correlation coefficieni.

The law (7) was obtained first by Sir Francis Oalton empirically, by studying observed frequencies. f As Pearson remarks:^ ‘That Galton should have evolved all this from his observations is to my mind one of the most noteworthy scientific discoveries arising from pure analysis of observations.’ Galton had not, at this stage, noticed that negative correlations exist, since he remarks: ‘Two variable organs are said to be correlated when the variation of one is accompanied on the average by more or less variation of the other, and in the same direction, ’§ and he speaks of correlation arising when two variations are the resultant of several causes, some common to both and some independent. The above analysis permits negative correlations. The more restricted one, however, is often valid and leads in particular to an account of intra- class correlation.

By integration w^e find

1 I x^\

P(dx\a,r,p,H) = (8)

t B.A. Report, Aberdeen, 3 885.

t Biometrika, 13, 1020, 25-45. This is a most interesting historical study. § Proc. Roy. Soc. 45, 1889, 135.

§2.5

DIRECT PROBABILITIES

73

Therefore

P{dy \a,T,p,x,H) ^

P{dxdy I O', T, p, H) P{dx I a, r, p, //)

That is, the probability of x is normally distributed with standard error or, and for given X the probability of y is normally distributed about prxla with standard error t^/(1— p^). The line y prir/o is known as the line of regression of // on x. Similarly the probability of y is normally distributed with standard error t, and that of x given y is normally distributed about x p<yyl'T, the line of regression of x on //. The lines of regression coincide only if p ^ 1 .

The expectations of x'^, y^, and xy, given a, p, r, H, are respectively CT^, par.

2.6. The characteristic function. Suppose that on a given law the chance of the variable x being less than an assigned value is f(x). Then the expectation of any function X{x) of a: is J A(a:) df{x) over the range of x; in which we must understand a Stieltjes integral if f(x) has discon- tinuities. These, if any, will all be positive jumps. The characteristic function Q(k:) is defined as the expectation of e^^, where k is })urely imaginary; thus

Je-d/(.r) (1)

and \0.{k)\ ^ 1. The integral is absolutely convergent because | df{x)

converges.

The integral

C + t'cO

r—loo

(X^ < Xg),

in which the path is a line parallel to the imaginary axis on the positive side, is equal to 1 if < x < x^, and zero if x < x^ or > X2, being the difference of two Heaviside unit functions. If we replace the path by the imaginary axis, except for a small semicircle about the origin, the integral is unaltered. Also the integral about the small semicircle tends to zero in the limit and the integrand is continuous. Hence we may replace the path by the imaginary axis, and

1 (Xi < X < Xo),

0 (x < Xj, X2 < x).

(2)

74

DIRECT PROBABILITIES

Chap. II

Now consider tlic sum

i'xi

Itti J k

-foo

over r, the ranges from to being so cliosen that all points of discontinuity of f(^) lie within them, and being some value between and On integrating witli regard to k, terms for not between

and X2 vanisli, while those bedween them contribute

1 {f(^r+l)-fi^r)} ->/(a^2)-/(a^l) (4)

in the limit when the intervals become indefinitely short. But the limit of the sum is by definition the Stieltjes integral

00 too i'oo

-A- j df{x) \ ~ dK == AL.-

2m J J j k

X=—co -ICO —ir>

by inverting the order of integration, which is easily shown to be valid.

When /(a:) is differentiable this leads to a case of Fourier’s integral theorem

dfjx)

dx

(6)

Similarly, if Cj, 62-* -j ^ variables whose chances are

independent and follow law^s given by shown

that the chance that

Xi ^2

ICO p 00

too 00-*

ico

^ id j (7)

ico

where the fl’s are the characteristic functions corresponding to the/’s. Hence the characteristic function of the sum of a set of variables following independent laws of chance is the product of their separate characteristic functions.

The characteristic function is intimately related to the expectations of the powers of x, where these exist. If we write

Mm = / x”^df(x).

(8)

§2.6

niRECT PROBABILITIES

75

we can call the mth moment of the law about the origin. If moments up to order m exist, we can differentiate (1) m times under the integral sign with regard to /c, and then for k = 0

fjm

Q(/c) = (9)

Thus, by Taylor’s tlu^orem,

i2(/c) 1 (10)

Z ! Tfl I

even though the complete Taylor series may not exist. For this reason il{K) is also called the moment -generating function. If we take the origin oi x at its exj)ectation, /x^ will be 0. Inspection of (1 ) shows that decreasing all values of :c by will multiply fl(/c) by and there-

fore if i2o(/c) is the characteristic function of x— /x^,

Qq{k) =r:r e~^f^^Q(K).

The coefficients of k^^/7iI in the expansion of logl2(/c) are called the semi-invariants or cumulants, when they exist, since the second and higher ones are independent of the origin and are additive for the sum of several variables. Also if y has a probability law^ g{y) such that g{y) ~ f(x) if y ax, where a is constant, the characteristic function of g{y) is

E{k) J dg{y) ~ J df{x) ~ D(a/c). (11)

The moment and the semi -invariant of g[y) of order m are times those of f{x).

If f2(K:) can be expanded in powers of /c, it will follow that the series represents an analytic function near #c = 0. But if any moment of the law diverges, the integral (1) defining Q.{k) will not exist for k on at least one side of the imaginary axis, however close to it, since the integral will contain a factor where c is real and not zero. Thus the integral will define a function only for purely imaginary values of k. It may be the value on the imaginary axis of some function analytic in the half-plane, but such a function, if it exists, w ill not be given off the axis by the integral. This applies to laws of Pearson’s Types IV, VII, and VI.

The integral may exist for all real k; this applies in all cases where the law has a finite range, such as the binomial and Type I laws. It is also true for the normal law. In that case the integral will exist for all K and be uniformly convergent in any bounded region of the k plane. It can therefore be integrated under the integral sign about any contour

76

DIRECT PROBABILITIES

Chap. II

in the k plane, and this integral will be 0 since [ (Ik ~ 0. Hence

c

by Morera’s theorem^ is an analytic function within any contour in the k plane, and must therefore be an integral function. J Tlien Q(k:) is expansible in powers of k over the entire plane.

There are cases where the integral exists for some complex values of K and not for others; for instance, the median law

(// -- I exp( \^'\/(i) (Lr/a.

Within the belt 1/a < J{{k) < 1/a the integral will define an analytic function. Outside this belt it diverges.

Thus we have two main types of ease. If all the moments of the law exist and the expectations of also exist, where c is some real quantity,

Q.(k) will be analytic near 0 and the coefiicient of will be iJL,Jn\ for all n. If moments up to order vi converge, but those of higher orders diverge, the integral does not define a function except for purely imagi- nary values of k. Its derivatives at k - 0 for imaginary k will give the moments correctly up to order but higher derivatives, if they exist, will not give the higher moments. We shall see that they do not neces- sarily exist.

2.61. The characteristic function is sometimes useful for actually calculating the moments. Thus consider the binomial law, according to which the chance of a sampling number less than I is

(1)

0

Then Q{k) ^ (2)

z=o

The coefficient of /c is which is therefore the expectation of 1. The moments about 7ix can then be derived by considering

{1 97 7? 1

7) 7hK^

= '^+-^^y + -^^yiy-^) + j,^xV+nxy(l-Gxy)}+..., (3)

wliere y = l—x-, whence the moments to order 4 about the mean are /i2 == /^3 = 7ixy{y—x); = Zn^xY^+nxy(l 6xy). (4)

t E. C. Titchmarsh, Theory of Functions, 1932, p. 82.

J I am indebted to Professor Littlewood for calling my attention to this point, in answer to a query.

§2.6 DIRECT PROBABILITIES

Pearson's parameters and are given by